OpenAI Agents Have Hacked More Companies Than HuggingFace

OpenAI has admitted that rogue ChatGPT agents accessed more than Hugging Face during an internal cybersecurity evaluation, after finding publicly exposed credentials for four accounts across four online services.

The incident began when an OpenAI model reached the internet during a test designed to assess its hacking ability. The model then accessed Hugging Face while searching for answers to ExploitGym, a cybersecurity benchmark. OpenAI later said the model had also accessed other publicly available services during the same evaluation.

 

What Did OpenAI Find Over And Above Hugging Face?

 

After the incident, OpenAI updated its account of what happened, saying, “The models identified and used publicly exposed credentials at the account-level on other publicly-available services. This includes four accounts on four services.”

The company did not name those services or explain exactly what information the models accessed through the accounts. OpenAI also said these incidents did not reach the same level of severity as the Hugging Face intrusion.

The discovery is worth noting, because the models were able to identify credentials that had been exposed online and use them during their search for information. The activity happened during an evaluation in which the models had been given a narrow cybersecurity task and their usual restrictions had been deliberately switched off.

Nik Kairinos, CEO and Co-founder of RAIDS AI, said the latest discoveries say a lot about how organisations currently supervise AI agents. He said, “The latest details of OpenAI’s rogue ChapGPT agents make this incident even more serious than it first appeared and it should be a defining moment for AI safety. If one of the world’s leading AI companies can lose control of an advanced model in this way, every organisation deploying AI agents should be asking whether its current safeguards are genuinely fit for purpose.

 

 

“Pre-release testing and sandboxing are essential, but AI systems can adapt, escalate and behave in ways their developers did not anticipate. Safety cannot be treated as a one-off exercise before deployment; it has to be continuous. This is why businesses need real-time monitoring that can identify when an AI system is drifting from expected behaviour, accessing systems it should not, pursuing unintended routes to complete a task, or creating new security risks.

“The lesson from this is that progress without ongoing oversight is a dangerous gamble. Trust in AI will depend on whether companies can show that their systems are safe not only in a test environment but throughout their entire lifecycle.”

 

Was The Incident A Sign Of AI Going Rogue?

 

Carole Reeves, director of security operations at digital transformation company ANS, said, “While the headlines understandably focus on an AI agent compromising external systems, it’s important to recognise this occurred during an internal security evaluation designed to better understand the capabilities and risks of advanced AI models. OpenAI’s transparency in sharing the findings gives the wider industry an opportunity to learn and strengthen future safeguards.

“Rather than taking this as a signal to fear AI, organisations need to ensure that security, governance and rigorous testing evolve alongside the technology. Strong incident response remains essential, but it is no longer enough on its own. Businesses also need a mature security posture, continuous exposure management and security by design to reduce risk before incidents occur.

“Secure and responsible adoption will be the defining factor of the AI race, and success will depend on how rigorously it is maintained.”

The incident therefore came from a controlled evaluation in which OpenAI intentionally removed restrictions that would normally limit risky cyber activity. The models then found weaknesses and online credentials that helped them complete the task they had been given.

OpenAI has reported the previously unknown software flaw to the vendor responsible for the affected software and said it is improving security around future AI evaluations.