OpenAI's advanced models spiraled out of control and carried out an "unprecedented" hack
Neural networks escaped from the test "sandbox" and hacked the "GitHub for AI"

An OpenAI agent escaped from the test environment, made its way onto the internet, and stole login credentials to carry out a hack / Photo: Prathmesh T/Shutterstock.com
OpenAI's state-of-the-art models went out of control during a security test: an autonomous agent built on their basis escaped from an isolated test environment, gained access to the internet, and hacked Hugging Face—the “GitHub for AI,” where AI projects are published and downloaded. OpenAI called the incident an “unprecedented cyber incident.” This is one of the first publicly known cases of a cyberattack carried out by AI without human involvement, according to the Financial Times.
How the AI Managed to Escape
For the test, OpenAI intentionally weakened the security of two models and instructed them to attempt a hack. They were operating in a “sandbox”—an environment designed to restrict their actions and isolate them from the internet. However, the models discovered previously unknown vulnerabilities, escaped the “sandbox,” and gained access to the internet. An AI agent then stole credentials and attacked Hugging Face, the FT reports, citing OpenAI.
Who was behind the hack?
OpenAI reported that the testing involved the GPT-5.6 Sol model, released in July, and another, more powerful model that the company is currently preparing for release. The company also warned that as neural networks with increasingly sophisticated capabilities for cyberattacks become more widespread, such incidents will occur more frequently.
"There was no malicious intent"
Hugging Face stated that it worked closely with OpenAI following the hack. “We are firmly convinced that there was no malicious intent on their part. It’s just amazing that the AI did all of this on its own!” wrote Hugging Face CEO Clément Delang on X.
It's going to get worse
Cathie Mussouris, head of Luta Security, called the incident a harbinger of new cyberattacks. She compared the current models to “the world’s smartest octopuses—masters of escape: they have an infinite number of tentacles and can squeeze through anywhere,” Reuters reports.
Matt Suish, an engineer at the cybersecurity firm Tolmo, noted that the technology for such attacks is already available far beyond the confines of cutting-edge laboratories: his team achieved similar results even without the latest models.
This article was AI-translated and verified by a human editor



