OpenAI said the autonomous models penetrated Hugging Face’s servers while searching for secret information that could help them pass an evaluation.
OpenAI, the company that makes the ChatGPT chatbot, has revealed that its artificial intelligence (AI) models autonomously breached another company’s production systems in what it called an “unprecedented cyber incident.”
The ChatGPT maker said in a July 21 blog post that the models—including its newly released GPT‑5.6 Sol and a more capable model still undergoing internal testing—compromised infrastructure operated by AI platform Hugging Face after escaping a restricted environment in which a cybersecurity evaluation was being conducted.
“The model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation,” OpenAI said.
In one case, the model assembled multiple attack vectors into a single cyber-strike, cobbling together stolen credentials with zero-day vulnerabilities on the Hugging Face servers before exploiting a remote code execution path to carry out a breach.
“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly,” OpenAI said.
The company released its preliminary findings less than a week after Hugging Face disclosed that an autonomous AI agent had penetrated part of its production infrastructure.
“This one was different from anything we had handled before in one important way: it was driven, end to end, by an autonomous AI agent system,” Hugging Face said in a July 16 blog post, adding that it initially did not know which model was behind the breach.
OpenAI said it discovered the anomalous activity internally before contacting Hugging Face, whose security team had already detected the model’s malicious behavior and were working on containing it.
“We are actively working with them to continue to investigate the incident,” OpenAI said, adding that it was “grateful for Hugging Face’s rapid and close collaboration on investigation and remediation.”
Although the models acted autonomously, the incident occurred in an evaluation setting in which OpenAI agents were pursuing advanced exploitation techniques, with some of the company’s normal cybersecurity safeguards intentionally disabled to measure their maximum capabilities.
“Autonomous, AI-driven offensive tooling is no longer theoretical,” Hugging Face said, adding that the key lesson from the breach is that defending online platforms requires not only heightened vigilance but also “using AI on defense to keep pace.”
OpenAI said it has imposed stricter controls on its research infrastructure while the vulnerabilities are patched, even though the measures could slow model development.
It also disclosed to developers the zero-day vulnerability that OpenAI’s models exploited in the breach, while strengthening cyber-protections in its testing environment.
By Tom Ozimek







