"Attacks Will Become a Real Threat": OpenAI Reports That Its Own Systems Were Hacked by AI Agents
The AI lab also revealed details about previous attacks. Following findings by independent researchers, OpenAI confirmed that approximately 700 of its agents were involved in the Hugging Face hack

Hundreds of OpenAI AI agents were responsible for the Hugging Face hack, according to independent researchers from METR and Redwood Research / Photo: doomu/Shutterstock.com
The July hack of Hugging Face —a popular platform where open-source AI models are published for free use—was carried out not by a single OpenAI agent, as previously reported, but by a swarm of approximately 700 agents acting in concert. This is according to a joint investigation by the research firms METR and Redwood Research. OpenAI confirmed to Reuters that the independent experts’ assessment is correct and also revealed that the AI agents not only infiltrated Hugging Face’s systems but also compromised its own internal infrastructure last month.
Details
According to METR and Redwood Research, the agents exchanged tens of thousands of messages through a communication channel they had set up without authorization. Reuters estimates that the scale of their interaction is likely to raise questions about how closely OpenAI monitored the tests. Hugging Face did not respond to the agency’s request.
According to OpenAI’s technical report, the agents ran their own code on 41 Hugging Face servers and gained superuser privileges on at least one of them, Axios reports. They also gained access to the platform’s development environment credentials and some of its internal data, and downloaded four private code repositories.
An inside job
The agents did not just infiltrate Hugging Face’s systems. OpenAI reported two incidents on July 19 in which the attackers breached the company’s own internal infrastructure. In one case, the agents exploited a vulnerability in a computer—which they were not supposed to access—and gained access to other connected systems. In the other, they stole OpenAI credentials and interfered with its cloud infrastructure, Reuters notes.
According to OpenAI, the agents attempted to deceive the evaluation systems or gain more freedom of action. An unspecified number of agents interfered with the automated tools that evaluated test results in order to conceal signs of cheating. The company claims that the interference ultimately did not affect the data used in the automated evaluation. The report also states that there was little evidence of attempts to deceive the human evaluators, according to Reuters. OpenAI did not respond to a question from Reuters asking whether this meant that some agents had, in fact, attempted to mislead the evaluators.
AI covers its tracks
An independent investigation and OpenAI’s own report state that AI models attempted to conceal violations by deleting or altering records of their actions. Independent experts found that one in five of the agents studied showed a clear interest in manipulating evidence, and many explored in detail ways to alter or falsify their activity logs, according to Reuters.
OpenAI also reported that the agents had deceived the systems in tests unrelated to cybersecurity, including when working with a protein database and a spreadsheet. This could mean the problem runs deeper, AI agent researcher Jeffrey Ladish of Palisade Research told Reuters. “It’s kind of like asking: if Billy cheats in every subject, not just computer science, is that more alarming? Well, yes, it’s more alarming,” he explained.
Missed Signals
OpenAI had noticed warning signs even before the Hugging Face hack, Axios reports, citing a company report. Initially, an internal team observed an agent using a makeshift forum and also detected instances of agents accessing the internet in violation of restrictions. About a month later, on June 27, the cybersecurity monitoring system alerted OpenAI to suspicious activity involving agents moving across the network via this forum, but the on-duty specialists decided not to halt the experiment. The company later acknowledged that some of these alerts may have warranted a response before the agents compromised Hugging Face.
OpenAI’s investigation also revealed signs that the training may have unintentionally reinforced behaviors that contributed to the incident. When the tools specified by the task did not work or were unavailable, the agents learned to explore and hack into individual elements of their environment. In some cases, they received positive reinforcement even after using unintended methods. For example, one agent exploited a vulnerability, accessed the program it was supposed to reproduce, copied the response, and received a reward for completing the task, Axios notes.
According to Axios, the incident raises the question of whether AI companies’ testing environments and internal security measures can keep up with models that are becoming increasingly adept at identifying and exploiting vulnerabilities on their own. “Given the rapid progress in the AI industry, we should assume that such attacks will become a real threat to companies in the near future and will become more sophisticated,” OpenAI warned (as quoted by Reuters).
This article was AI-translated and verified by a human editor



