"Three guys who follow Claude" hacked OpenAI's internal system using AI
The Financial Times refers to them as “white-hat hackers”—they reported the vulnerability to OpenAI, for which they received a reward from the company

Cybersecurity researchers used Anthropic's AI model, Claude, to gain access to a ChatGPT account. Photo: Koshiro K/Shutterstock
A small group of cybersecurity experts—so-called “white hat hackers” (or “ethical hackers,” as the Financial Times refers to them)—used Anthropic’s Claude software to gain access to an OpenAI employee’s ChatGPT account. This allowed them to view the company’s internal code repository—a specialized, closed repository where developers collaboratively write, test, and store the source code for algorithms and software—and propose changes to it. After discovering the vulnerability, the “white-hat hackers” submitted a report about it to OpenAI’s developers, for which they received a $6,500 reward from the AI lab, according to The Wall Street Journal (WSJ).
Details
The OpenAI hack began on July 23, when experts from a small company called Hacktron AI used Anthropic’s Claude software to exploit a vulnerability on the OpenAI community forum (which runs on the Discourse platform). Through this vulnerability, the researchers gained access to internal authentication systems and took control of a ChatGPT account belonging to one of the employees, according to the Financial Times (FT). This account had direct access to a private GitHub repository. Using ChatGPT as an interface, the researchers were able to view corporate files.
According to them, they stopped the attack as soon as they realized they had accessed confidential information. But before submitting their report to OpenAI, the researchers managed to request some changes. They instructed the algorithm to add the string “Hacktron AI Team PoC” and links to their X accounts to the documentation file. This served as proof that they had indeed gained access to the company’s confidential data.
OpenAI confirmed that an internal audit of GitHub revealed only “limited read access” to the metadata of a private repository and its code change history. The issues have now been resolved. Discourse representatives reported that they patched the security vulnerability on July 25, the day they received notification.
Researchers from Hacktron AI received $6,500 from OpenAI as part of a bug bounty program—an initiative designed to identify vulnerabilities before malicious actors can exploit them, according to the Financial Times. “We thank the researchers for contacting us and sharing their findings. We have restricted access rights for Community login tokens and revoked the affected tokens and sessions,” OpenAI said in a statement.
Anthropic stated that the disclosed data makes it possible to “understand how close the world is to achieving recursive self-improvement”—the point at which AI will be able to train and improve itself or new models. This threshold lies at the heart of concerns that AI systems will become harder to oversee, leading to a loss of human control.
Anthropic added, however, that their models do not yet operate completely autonomously. In 90% of tasks, the AI “collaborates” with humans and performs a large volume of work under their supervision.
Why Is This Important?
Against the backdrop of the AI race between the U.S. and China, the researchers who hacked OpenAI note that this attack proves that advanced hacker groups have a very real chance of gaining access to the country’s technological secrets, the WSJ reports. “I don’t think we’re as strong as Chinese attackers,” said Mohan Pedhapathi, CTO of Hacktron AI. “We’re just three guys with subscriptions to Claude and Codex (an AI model for programming—ed. note from Oninvest).”
This incident highlights the challenge of protecting corporate secrets in the age of AI hacking, according to Joshua Saks, CTO of Abundant Security (which specializes in AI security), who reviewed the Hacktron AI report. “The world is full of security flaws in software. The reason we haven’t found them all is that, until last year, there were only a few thousand people who were experts at finding these flaws,” he said. Now, AI agents are making this capability accessible even to less-skilled professionals, Saks added.
Context
The hack occurred just two weeks after a swarm of more than 1,000 OpenAI AI agents escaped from a test environment and attacked the startup Hugging Face. This incident was followed by a series of events in which AI agents and models from various leading AI labs in the U.S. and China—including Anthropic, OpenAI, Meta, and Moonshot—spun out of control.
Against this backdrop, on September 12, Antrophic CEO Dario Amodei called for a slowdown in the development of AI technologies, arguing that progress is moving too quickly and developers are unable to keep up with ensuring safety and minimizing potential harm. OpenAI CEO Sam Altman and SpaceX CEO Elon Musk agreed with Amodei.
This article was AI-translated and verified by a human editor



