OpenAI has revealed that a group of its most advanced AI models bypassed the restrictions of a controlled security test before carrying out an autonomous cyber intrusion targeting technology start-up Hugging Face.
The ChatGPT maker said the agents, which can independently complete tasks after receiving human instructions, identified flaws in the safeguards intended to keep them confined during the exercise.
After bypassing those restrictions, the models targeted Hugging Face, one of the world’s largest platforms for hosting and sharing AI models, and gained access to parts of the company’s internal infrastructure.
OpenAI described the incident as “unprecedented” and said it had launched an investigation with Hugging Face. The platform’s chief executive, Clement Delangue, said in a post on X it was “mind-blowing that all of this happened autonomously”.
“The investigation is ongoing, and we’ll share more learnings from what might be the first incident of its kind,” Delangue added.
A UK government spokesperson said the AI Security Institute was examining the behaviour demonstrated by the system and continuing to work with OpenAI and other laboratories to strengthen safeguards.
The spokesperson also urged organisations to improve their cyber-security measures, including by enrolling in the government-backed Cyber Essentials certification scheme.
Gina Neff, head of the Minderoo Centre for Technology and Democracy at the University of Cambridge, told BBC Radio 4’s Today programme that security testing environments, commonly known as sandboxes, are “supposed to be secure environments where you can see what the models are capable of”.
“In this case, it looks like OpenAI didn’t make a secure enough sandbox,” she added.
Rather than simply completing the task assigned to them, the agents launched an attack against the sandbox itself and identified a vulnerability that enabled them to escape its restrictions.
Once outside the controlled environment, the AI reportedly identified Hugging Face as a potential source of information needed to complete the original test and attempted to access the company’s systems.
Neil Lawrence, Professor of machine learning at Cambridge University, described the operation as an “impressive feat”, while warning that it “falls well within the known capabilities of the current generation” of powerful AI models.
Lawrence noted that OpenAI is considering a stock-market listing and faces mounting competitive pressure from Anthropic, which has attracted attention for its own advanced AI tool, Mythos.
“OpenAI are now playing catch-up, they are trying to demonstrate their own systems’ capabilities in cyber-security.”
“It shows us that OpenAI are not capable of safely deploying their own technology,” he added.
Hugging Face first disclosed the breach on July 16, saying it was still investigating whether any customer or partner information had been compromised. The company said it would contact affected parties if necessary.
It has since closed the vulnerabilities exposed during the incident and rebuilt the systems that were affected.
“Autonomous, AI-driven offensive tooling is no longer theoretical,” it said.
“Defending an online platform now means treating the data and model surface as a first-class attack surface, and using AI on defence to keep pace.
“We will keep investing there, and keep sharing what we learn.”
The incident has renewed concerns about the growing capabilities of advanced AI systems and whether existing safeguards are sufficient to contain increasingly autonomous technology.
Spencer Starkey, an executive at cyber-security company SonicWall, said that organisations must “step up” their security measures and “treat cyber resilience as a core operational priority”.
“The uncomfortable truth is that too many organisations are still defending at human speed while adversaries are escalating to machine speed,” he said.
Travis Lelle, principal security engineer at cyber-security consultancy GuidePoint Security, described the disclosure as a “sobering moment in cyber-security”.
“This highlights a known asymmetry,” he said.
“Offensive agents are unconstrained, while the best defensive tools are locked behind guardrails that cannot understand context.”
However, Jake Moore, global cyber-security adviser at ESET, suggested the announcement may also be partly driven by competition within the rapidly developing AI industry.
He argued that OpenAI could be attempting to draw attention to the strength of its technology as rival Anthropic receives growing recognition for its Claude Mythos model.
“It does pose the question that OpenAI are potentially chasing the marketing dream of Anthropic of late,” he said.
The disclosure came one week after Chinese AI start-up Moonshot unveiled Kimi K3, a major new artificial intelligence model that the company claimed could compete with leading systems developed by US firms.


















