Meta AI Escape: Model Hacks Third-Party Service During Test

It feels like every day a new AI company comes forward with news that its AI model was able to escape during testing and hack into another organisation’s systems. In the past few weeks alone, OpenAI, Anthropic and the UK’s AI Security Institute (AISI) all publicly revealed to have suffered one of these incidents. Yesterday, Meta joined the list, making it the fourth recent incident of its kind.

As Damian Skeeles, Senior Solution Engineer Manager at Filigran, succinctly puts it: “It appears Frontier models are getting FOMO now, and each needs to have their own lab breakout story.”

Naturally, these incidents have raised significant alarms amongst cybersecurity professionals. Oliver Simonnet, Lead Cybersecurity Researcher at CultureAI, said, “this latest report that now Meta’s AI model also compromised an external organisation during testing is now exhausting as well as alarming.”

In the case of Meta, a spokesperson told the BBC that the hack had been caused by a “misconfiguration” by its independent tester. Meta said the security trials were conducted by Irregular, the same AI security vendor that tested Anthropic’s AI model after it gained access to the systems of three other companies.

This differs from the more sophisticated sandbox escape we saw in the case of OpenAI. Due to the less sophisticated nature of the Meta hack, Simonnet describes it as “even more unacceptable.”

Simonnet continued: “If a third-party organisation outside of an agreed scope was compromised during an automated penetration test, this would trigger immediate legal, regulatory, and industry scrutiny. The fact that it happened during a ‘test’ would in no way excuse the failure to contain the activity.”

“No third party should become an involuntary participant in an AI capability test, let alone multiple times in the space of a few days,” Simonnet concluded.

There’s also concern that cybercriminals are using similar tools for AI modelling and testing too. Cybercriminals using the same tech as security teams is a story as old as time. If new tech becomes available, black hat hackers will race to use it, in parallel (if not ahead of) legitimate security teams.

On this, Paul Bischoff, Consumer Privacy Advocate at Comparitech, said: “The tools used by black hats are often the same ones used by white hats. Long before AI, cybersecurity tools like Metasploit, Cobalt Strike, Censys, Shodan, and Wireshark were widely used by both cybercriminals and the people tasked with stopping them.”

“Businesses use these tools to test their own software and networks for vulnerabilities, which they then patch. Cybercriminals use these tools to find vulnerabilities that haven’t been patched. AI is the same; it’s agnostic.”

“Although the flagship models used by most people might have some guardrails in place to prevent abuse, open-weight models have no such restrictions. Hackers and cybersecurity engineers will both use open-weight models to find vulnerabilities,” Biscoff continued.

It is likely that similar incidents will occur. However, time and analysis is necessary to help us learn from these incidents and understand how and why they have happened. As Skeeles of Filigran said in the case of Meta: “We’ll have to wait for the technical analysis to see if there were any novel techniques to worry about” and, in the race for AI security, that analysis can’t come soon enough.