HomeNews
Share

Anthropic's AI models hacked into the systems of three organizations during testing

One of the models continued the attack even after realizing it was connected to the internet

Albert Fahrutdinov

Albert Fahrutdinov

reporter Oninvest
Anthropics models hacked real-world systems during testing / Photo: Photo For Everything/Shutterstock.com

Anthropic's models hacked real-world systems during testing / Photo: Photo For Everything/Shutterstock.com

AI models from Anthropic’s Claude series hacked into the live systems of three organizations during a pre-launch cybersecurity test, the startup reported on its blog. The incidents involved Opus 4.7, released in April, and two of Anthropic’s most powerful models—Mythos 5 and an internal research model that was not intended for release.

The tests were conducted in a “capture the flag” format: the models were tasked with hacking into other systems to find hidden information. The older version continued the attack even after realizing it was connected to the internet. The newest version stopped once it realized it was online.

How did this become possible?

Due to a misunderstanding between Anthropic and its external partner, Irregular, the test environment remained connected to the internet, even though Claude was supposed to operate in a simulation without internet access. As a result, the models mistook real-world systems for part of the exercise. Unlike in the OpenAI case, they did not exploit any zero-day vulnerabilities (hidden bugs that developers are not yet aware of): internet access was enabled by the test configuration, according to Axios.

OpenAI CEO Sam Altman / Photo: Antonello Marangi/Shutterstock.com

OpenAI Models Have Hacked a Second Company: What Will the Consequences Be?

The tests were conducted without the security measures included in publicly available versions, so that researchers could assess the AI’s capabilities. According to Anthropic, if those restrictions had been in place, they would have blocked the hacking attempt. Claude used simple techniques: it exploited weak passwords and interfaces that did not require authentication.

Three Hacks

Following OpenAI’s report of unauthorized leaks of its models onto the public internet, Anthropic reviewed more than 141,000 of its test runs and discovered three similar incidents. The first of these occurred in April. In one instance, Opus 4.7 was unable to reach a fictional target but found a real website with the same name and hacked it.

Believing it was operating in a simulation, Mythos 5 created and uploaded a malicious Python package to the public PyPI repository. It remained available for about an hour. During that time, it was downloaded and executed on 15 real systems. One of them belonged to a cybersecurity company that automatically installs Python programs to detect malicious software. As a result, the Claude code was able to steal credentials and gain access to other infrastructure assets.

In the third case, the internal research model failed to reach the fictitious target, scanned approximately 9,000 targets, and hacked into a company’s application that was accessible via the internet. It then realized that it had ended up in a cloud account “completely unrelated to the capture-the-flag challenge” and halted the attack.

Why Is This Important?

The incidents involving OpenAI and Anthropic indicate that the models remained focused on performing the tasks assigned to them, rather than pursuing their own goals, Axios reports.

Anthropic disclosed the breaches nearly four months after announcing Mythos, the distribution of which the company had strictly limited due to the system’s power and the potential risks posed by the project, Bloomberg reports.

"Safety testing is conducted before a model is released precisely because we don't yet know what it's capable of," Anthropic said.

This article was AI-translated and verified by a human editor

Share

Trending

Stock Screener
Buy
Sell


















Small Caps
Investment and Finance News