We keep hearing that AI agents are doing what we don’t want them to. Now OpenAI says it’s holding back parts of the rollout of its latest model, GPT-6 Astra, while strengthening safety and alignment safeguards after testing identified potentially concerning autonomous and cybersecurity capabilities.
But there’s something of a contradiction here that should make us more than a little wary of assuming we’re being told everything we need to know. So warns the CEO of Doxis, Dr John Bates, a former Cambridge University computer science academic turned serial software company founder and enterprise AI leader.
“In the case of the Australian attack, an agent behaving in unexpected ways isn’t new, but it does further the case that some sort of guardrails do need to be put in place around these ever-improving autonomous AI actors,” he says, adding: “What’s next: agents cracking a national intelligence agency’s servers? Things seem to be moving at such a pace that we need to put better guardrails around these systems.”
Bates, who also serves as a non-executive director at Sage, argues that we shouldn’t be surprised when an agent behaves in unexpected ways. That isn’t new, it’s rather what you might expect from an agent.
We Can Handle Smart Responses But Unpredictable Is Another Matter
His argument is that if agents going off-script and invading other systems isn’t what we expected, then we need governance around it. But aren’t these systems, in some sense, simply doing what they were asked to do?
“If the behaviour wasn’t what was expected, then we need governance around it. But there’s an interesting question here: are these systems not actually doing what we asked? The Hugging Face incident and then this hack were the result of an internal cybersecurity evaluation by OpenAI: isn’t this exactly the kind of initiative behaviour we should be expecting, and indeed praising?”
For Bates, the new factor in cybersecurity and AI risk isn’t that systems are smart. We’ve had computer systems for decades, from algorithmic trading to real-time operating systems, that move faster than we can.
More from Artificial Intelligence
- A Court Just Gave Human Voices Legal Protection From AI – But How Far Does That Go?
- What Does A Modern SME Really Need From AI?
- SOFTSWISS Launches 2027 iGaming Trends Report: An Industry At The Crossroads Of AI, Growth And Regulation
- Why Teaching AI Anatomy Is The Hardest Leap In Digital Animation
- Do Domain Names Still Matter In An Age Of AI-First Internet?
- The AI Confession Trap: Why Your Chat History Has No Legal Privilege In Court
- Salmon Introduces Execution Verification Infrastructure (EVI) For Securing AI Agents And Autonomous Systems
- 97% Off Claude And GPT: Why Stolen AI Subscriptions Are Flooding The Dark Web
What has changed is that, by design, large language models are non-deterministic, so even if you tell an AI agent it can do X but not Y, it may still behave outside those boundaries.
That’s because models may have learned lots of new things today that means the answer they gave yesterday won’t be the same: “Some might argue that we’ve always had smart systems, so why now but a major factor that’s changed is that this class of software is a non-deterministic system. That means that even if you tell an AI agent it can do X but not Y, it may still behave outside those boundaries.”
He adds, “It’s a bit like chaos testing, where you deliberately unleash unexpected conditions on a system to see how it responds except here, we have something much more sophisticated: a naughty child with Dr House-level intelligence trying to find a way to break your system.”
Hand-wringing over AI failing to do what we told it to is no defence, he concludes. Instead, he argues, we should combine AI with traditional, rules-based technologies, putting controls around it to enforce clear boundaries and keep it within them: “There is a strong case for combining AI, which acts in probabilistic ways, with traditional, rules-based technologies, and bringing in controls to enforce boundaries and keep AI within them.”
This would anchor AI in mainstream business applications, keeping it grounded in the real world. He believes this kind of practical lock-in could deliver greater long-term benefits to society than political intervention:
“We already have an example of regulation in the EU AI Act, which contains many provisions around things AI shouldn’t be used for. US technology providers may decide to restrict certain technologies, or not allow them to be used in particular applications, because they don’t want to risk falling foul of European regulation.
“It comes from a good place, and regulation is important but my concern is that Brussels has gone too far in places, and some of the people writing the rules don’t fully understand the technology they are trying to regulate.”
