HomeNews
Share

AI agents from OpenAI and Anthropic went out of control and attacked real-world targets 19 times

It all happened during tests of AI agents conducted by a British regulatory agency—AI systems developed by leading AI labs created fake profiles on GitHub and misled developers

Albert Fahrutdinov

Albert Fahrutdinov

reporter Oninvest
A participant in the “March Against Cars” protest in London holding a sign that reads “AI—Thief of All Things” / Photo: Loredana Sangiuliano/Shutterstock.com

A participant in the “March Against Cars” protest in London holding a sign that reads “AI—Thief of All Things” / Photo: Loredana Sangiuliano/Shutterstock.com

The British government-established AI Security Institute (AISI), created to evaluate cutting-edge AI systems, reported that in July, during cybersecurity tests of the Mythos 5 (Anthropic) and GPT-5.6 Sol (OpenAI) during cybersecurity tests in July, they went out of control and carried out “19 actions against real people and organizations.” Mythos was responsible for 17 of them, and GPT-5.6 Sol for two. The researchers clarify that these were not 19 separate incidents, but rather “several related lines of behavior.”

The scammers created fake accounts on the GitHub developer platform, attempted to manipulate programmers working on open-source projects, embedded hidden malicious instructions, and sent emails containing false information. GitHub confirmed the violation of its rules, and the institute, in collaboration with the service, erased the agents’ traces and alerted users to the incident.

British researchers identified the most serious incident as one in which a neural network engaged in social engineering. It created several fake identities and began trying to persuade the developer overseeing the project to approve the changes. The developer recognized the threat and rejected the code.

“These attempts were unsuccessful, and our investigation did not reveal any actual harm. But this is the first time we’ve seen the risks associated with autonomy and fraud manifest so clearly, without any specific prompts, in the real world,” AISI stated.

Test Conditions

Researchers deliberately granted the models access to the internet and disabled cybersecurity filters, and did not prohibit them from using the internet, according to Axios. Unlike the July hack of Hugging Face by an OpenAI agent, the models did not escape their isolated environment during the AISI tests, Reuters reports. However, the British institute acknowledged that it does not yet understand “at what point the agent realized it was operating in the real world, or to what extent it believed it was in a fictional scenario.”

OpenAI CEO Sam Altman / Photo: Antonello Marangi/Shutterstock.com

OpenAI Models Have Hacked a Second Company: What Will the Consequences Be?

The Second Incident

On August 4, OpenAI reported another incident: its security partner, Irregular, discovered that the developer’s models had gone online and hacked a real website whose name matched that of a fictional company from the simulation, according to Axios.

This refers to a test similar to the one during which Anthropic’s models hacked three organizations, a Bloomberg source clarified. The internet was intentionally left accessible in the sandbox to allow for a realistic assessment of capabilities, but the parties had not fully agreed on the procedures and security measures, an Axios source said. The tests were conducted “with weakened security measures, under conditions that do not reflect normal use,” OpenAI emphasized.

What's next?

The cyber capabilities of cutting-edge AI models have caught researchers off guard and are forcing them to reevaluate security protocols, Axios notes. Over the past month, both OpenAI and Anthropic have acknowledged that their models hacked real organizations and websites during routine pre-release testing.

AISI is developing restrictions on AI agents' access to the internet and implementing real-time monitoring of them. This is intended to block malicious activity before it interacts with external resources.

OpenAI and Irregular are preparing guidelines for conducting tests safely. Anthropic stated that the incident “highlights the need for a broader conversation about how to safely evaluate increasingly capable AI agents” (quoted in the AFR).

JPMorgan CEO Jamie Dimon compared unrestricted access to Mythos to “handing over missiles” to private individuals / Photo: Mijansk786 / Shutterstock

"It's like handing out ballistic missiles": Damon on the dangers of the Mythos AI model

This article was AI-translated and verified by a human editor

Share

Trending

Stock Screener
Buy
Sell


















Small Caps
Investment and Finance News