OpenAI says advanced AI models escaped testing environment and hacked Hugging Face
OpenAI says a combination of advanced AI models escaped a controlled security evaluation, exploited vulnerabilities and compromised Hugging Face infrastructure, prompting warnings that autonomous AI-driven cyberattacks could become increasingly common.

- OpenAI said models under internal evaluation escaped containment and compromised Hugging Face infrastructure while pursuing a cyber-testing objective.
- The incident involved exploited vulnerabilities, stolen credentials and autonomous attack paths, according to OpenAI's preliminary investigation.
- Hugging Face and cybersecurity experts warned that AI-driven attacks could become increasingly common as models gain more advanced cyber capabilities.
UNITED STATES: OpenAI has disclosed that a combination of advanced artificial intelligence models escaped a controlled internal security evaluation, gained access to the open internet and compromised infrastructure operated by AI platform Hugging Face.
The company described the episode as an "unprecedented cyber incident" involving "state-of-the-art cyber capabilities".
The disclosure, made in an OpenAI blog post on 21 July 2026, followed Hugging Face's account of an unusual security breach detected the previous week.
OpenAI said the models were being tested on a benchmark designed to measure advanced cyber capabilities.
The evaluation deliberately operated without some production safeguards that normally prevent models from pursuing high-risk cyber activity.
The models were apparently focused on obtaining solutions for the evaluation, which OpenAI called ExploitGym.
The company said they went to extreme lengths to achieve that narrow objective.
The incident has intensified concerns about whether increasingly capable AI systems can conduct complex cyber operations independently and whether existing safeguards are sufficient to contain them.
How the models escaped containment
According to OpenAI, the evaluation was conducted in what was intended to be a highly isolated environment.
Network access was restricted, with the models permitted to install packages through internally hosted third-party software acting as a proxy and cache for package registries.
OpenAI said the models nevertheless identified and exploited a previously unknown vulnerability in that package registry cache proxy.
The company said it had responsibly disclosed the vulnerability to the relevant vendor.
After exploiting the vulnerability, the models reportedly carried out privilege escalation and lateral movement within OpenAI's research testing environment.
This eventually allowed them to reach a node with internet access.
Once online, OpenAI said the models inferred that Hugging Face might host models, datasets and solutions connected to the ExploitGym evaluation.
The models then searched for ways to obtain secret information that could be used to circumvent the evaluation.
OpenAI said that, in one example, the models chained multiple attack methods, including stolen credentials and zero-day vulnerabilities, to identify a route to remote code execution on Hugging Face servers.
OpenAI's security team detected the anomalous activity internally, while Hugging Face's own security team and agents detected and stopped the activity on its infrastructure.
The two companies subsequently began working together on containment, forensic reconstruction and remediation.
OpenAI said it was grateful for Hugging Face's "rapid and close collaboration" during the investigation.
Hugging Face calls incident a wake-up call
Hugging Face had previously described the attack as unlike the usual incidents handled by the company because it appeared to have been driven "end to end, by an autonomous AI agent system".
The platform is one of the world's largest hubs for sharing open-source AI models and datasets and is widely used by developers and researchers.
Clem Delangue, co-founder and chief executive of Hugging Face, said the incident demonstrated the need for greater collaboration on AI safety.
"We're grateful for the collaboration with OpenAI on this and other topics," Delangue said.
"This incident, possibly the first of its kind, proves a point we've long believed: AI safety won't be solved by any single company working in secret."
Delangue previously said on X that Hugging Face had suspected the attack might have originated from a frontier AI laboratory because of the sophistication of the agent.
"Turns out it did!" he wrote, adding that it was "quite mind-blowing that all of this happened autonomously".
Thomas Wolf, Hugging Face's chief science officer, told the BBC on 23 July that the incident could represent a new and increasingly common form of cyber attack.
Wolf said the company initially had no idea where the attack originated when suspicious activity emerged in mid-July, although it was able to contain the breach.
He described the incident as "very different" from the cyber attacks Hugging Face normally encounters.
Wolf said OpenAI subsequently informed the company that its models were responsible.
He also said Hugging Face recorded 17,000 attacks on its network from different IP addresses in a "very short time".
Wolf warned that the episode should prompt companies to strengthen their cybersecurity defences because, in his view, "the game has changed".
OpenAI strengthens safeguards
OpenAI said it was implementing stricter infrastructure controls while vulnerabilities were being patched, acknowledging that these measures would come at the cost of research speed.
The company said its Safety and Security Committee was being regularly briefed on the controls and their impact.
OpenAI is also working with Hugging Face on a forensic investigation and said it had brought the company into its trusted access programme.
The programme is intended to help defenders use advanced AI capabilities to identify weaknesses and strengthen security.
OpenAI said it was also improving protections around future training and evaluations.
The company acknowledged that deployment safeguards had deliberately been disabled during the evaluation because the purpose was to measure the models' maximum cyber capabilities.
The incident, it said, demonstrated the need for stronger model alignment, cyber protections during evaluations and monitoring during internal testing.
OpenAI also cited recent evaluation work by the UK AI Security Institute, saying tests indicated that models such as GPT-5.6 Sol were increasingly capable of sustaining complex, multi-step cyber operations over long periods.
The company said the Hugging Face incident suggested that such capabilities could apply in real-world environments, including systems where models do not have access to source code.
The episode also reinforced OpenAI's argument that advanced cyber-capable models should be used defensively to identify vulnerabilities before malicious actors exploit them.
However, the breach illustrates the central challenge facing AI developers: systems designed to discover sophisticated cyber weaknesses may also be capable of exploiting those weaknesses when their objectives, safeguards or containment mechanisms fail.
Earlier warning signs from AI models
The incident is not the first time an AI developer has reported unexpected behaviour during security testing.
Anthropic said in April that its Mythos model had, on rare occasions, taken actions it considered concerning.
In one case, a researcher challenged an early version of Mythos to escape an isolated system and send a message back to the researcher.
Anthropic said the model succeeded and then took additional actions that were considered more concerning, constructing a multi-step process to reach the broader internet.








