One expert compared the OpenAI model hack to a dystopian thought experiment in which a machine eventually destroys humanity.
What if an artificial intelligence model is given a task and uses every conceivable resource at its disposal to complete it, even to the detriment of humanity itself?
That is what some are now fearing after an OpenAI model broke out of a testing sandbox and used zero-day exploits to hack into Hugging Face, an open-source community for AI and machine learning, to crack a problem it was instructed to solve.
“This is some of the clearest evidence yet that an AI model can run a complete cyberattack from start to finish without a human steering it,” Andrew Jones, cofounder and chief product officer at cybersecurity firm Adaptive Security, told The Epoch Times.
OpenAI, developer of the popular large language model-powered ChatGPT chatbot, acknowledged the security breach on July 21, stating that two of its advanced models escaped a restricted testing environment and broke into Hugging Face’s software infrastructure.
“After investigating, we now know that this particular incident was driven by a combination of OpenAI models—including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes—while being internally tested on a benchmark of cyber capabilities,” OpenAI stated at the time.
Some analysts referred to the incident as an example of AI “going rogue,” “scheming,” or pursuing goals that were at odds with its human testers.
However, multiple AI and cybersecurity experts interviewed by The Epoch Times called this depiction misleading, arguing that the models were pursuing OpenAI’s stated objectives within a testing sandbox in which the company had relaxed some of its usual safety constraints.
“‘Scheming’ implies the model wanted something other than what we asked for. It didn’t. Every step was in service of the goal we set,” AI expert Anik Devaughn told The Epoch Times.
Devaughn, who has worked in the industry for years and founded AI firms Wired to Create and Karo, said the Hugging Face breach is more concerning than AI “scheming.”
“A machine with hidden motives is a problem you can look for,” he said. “A machine with no motives at all, executing your instructions past the point you stopped imagining, is a problem you have to engineer against.”
One expert called it a “canary in the coal mine” situation.
“If this was a human black hat hacker doing it, there would be arrests and litigation, and you name it. It‘d be illegal. It’d be a cybercrime,” Nicholas Nadeau, founder of Onix AI, told The Epoch Times.
“This was a full-on agentic automated breach by an entity, and so we start to ask the questions of who’s responsible. Is it OpenAI? Does the agent, the AI, have a certain level of responsibility? Do the researchers?
“We don’t even have rules, regulations, laws, or even a thought process or framework around defining what responsibility means.”
Relaxed Testing Environment
OpenAI said it was evaluating its models in a sandbox it calls ExploitGym, which has relaxed security constraints for testing.
“We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity,” the firm said, describing the relaxed safety protocols that the model exploited.
The firm asked the models to “pursue advanced exploitation using complex attack paths” to determine their cybercapabilities, but said the environment was “highly isolated” and had internet access limited to the models’ ability to “install packages through an internally hosted third-party software.”
When the models were set to work, they found and chained multiple security exploits in both OpenAI’s research environment and Hugging Face’s production infrastructure to locate the test’s solutions within the latter firm’s databases.
Juan Pedro Márquez, a cloud solution architect specializing in AI and agents at Microsoft’s Europe, Middle East, and Africa division, said the “unsettling part” of the incident was not the breach itself, but how the model achieved that task.
“OpenAI lowered its safety refusals for a benchmark test, and the model reasoned that Hugging Face’s infrastructure probably held the answer key, broke out of its sandbox, found a real zero-day, and used it,” Márquez said. “Nobody told it to.”
Aimee Simpson, director of product marketing at cybersecurity company Huntress, also cast doubt on the narrative that the models were scheming when they broke out of OpenAI’s testing environment.
“It’s not like the system itself was malicious,” she told The Epoch Times.
“The developers still had the objective of acquiring that information; they just didn’t think it would be able to bypass the limits of its test,” she added, describing the path the model chose as “highly creative and extremely unusual.”
By Jacob Burg







