A new report said around 700 OpenAI agents were a part of the Hugging Face attack last month, and 1,200 collaborated through a message board.
About 700 OpenAI agents were involved in the breach on AI startup Hugging Face, independent researchers said in an Aug. 26 report.
OpenAI revealed last month that its AI agents—AI programs that run with minimal human intervention—were the ones that hacked into Hugging Face’s system.
Wednesday’s report by METR and Redwood Research revealed that hundreds of agents collaborated on the attack—about 700 out of 1,200 agents that sent 70,000 messages to each other between June 26 and July 13 on an unsanctioned messaging board.
“This incident is the first known case of an automated agent collective acting offensively without authorization,” OpenAI said in a separate investigative report that it also published on Aug. 26.
The ChatGPT-maker said it responded to the incident with the assistance of external advisers, including cybersecurity giant CrowdStrike. OpenAI’s report didn’t specify how many agents were involved in any of the attacks, but mentioned that agents hacked into its own infrastructure as well during evaluation tests.
The company said agents escaped their testing environment and accessed other connected systems while other agents obtained OpenAI credentials and potentially exposed its research systems to internet traffic.
The San Francisco-based company added that on June 27, a monitoring tool alerted it of unusual activity in its research environment, which responders linked to an agent using an improvised message board, advising that stopping the evaluation was not necessary.
OpenAI noted that its standard safeguards for user models were turned off during these tests.
The Epoch Times reached out to OpenAI for further comment.
OpenAI said it conducted routine testing for an unreleased research model and GPT-5.6 Sol beginning in May, where its agents attempted to solve cybersecurity tests in an isolated environment separate from OpenAI’s internal infrastructure and the internet.
The agents found ways to communicate and exchange information through a message board on a third-party service that provides agents with controlled access to software packages.
The company said that on July 8, agents exploited a vulnerability to gain internet access and between July 11 and July 13 compromised Hugging Face’s infrastructure, which they then used while trying to solve the evaluation tests.
“The model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation,” OpenAI said in a July 21 statement.
Hugging Face identified the breach but, at the time, was unaware who was behind the agents.
“This [intrusion] was different from anything we had handled before in one important way: it was driven, end to end, by an autonomous AI agent system,” the company said on July 16.
OpenAI said it detected suspicious activity on July 19, and a day later determined these models were likely related to the Hugging Face breach and informed Hugging Face. On July 21, OpenAI publicly revealed its role in a blog post.
By Dylan Morgan






