Secret AI Chats: OpenAI Agents Communicated on Dozens of Websites, Circumventing Restrictions
Reuters compares this to the behavior of students who are not allowed to talk during an exam: they pass notes to each other in the restroom stall

OpenAI agents secretly communicated on dozens of websites, circumventing restrictions. Photo: Stock all/Shutterstock
OpenAI’s AI agents, which had spun out of control, used more than ten websites for unauthorized communication. This is evidenced by data from six independent research groups and information reviewed by Reuters. They revealed that the agents’ unauthorized activity—which the company had concealed for several months—turned out to be significantly more widespread than initially reported in early September. At that time, only a single hack of a German platform for programmers had come to light.
What Happened
On September 4, researchers reported that a group of OpenAI AI agents had taken over the German-language website DseWiki, which is geared toward the developer community, and turned it into a makeshift platform for messaging and cheating on tests. This site is one of the so-called wiki platforms that allow for collaborative content editing, similar to “Wikipedia.” OpenAI kept the incident under wraps while dealing with the fallout from the July hack of the Hugging Face open-source repository, Reuters reports. Now, independent analysts claim to have discovered several more websites where the same agents left similar messages.
The researchers identified activity by comparing data rows from a German website with identical code on other resources. They also analyzed matching usernames or highly specialized search queries—for example, those regarding cancer incidence statistics in Iowa. In a number of cases, the traces led to IP addresses belonging to Microsoft Azure infrastructure, which OpenAI sometimes uses.
The exact number of affected websites varies, but all experts interviewed by Reuters agree that there were more than ten such sites. These are primarily collaboratively edited wiki sites, text-hosting platforms, and link-shortening services operated by two universities.
Most of these sites were little-known. Traces of AI agents were found, for example, on a wiki site for in-depth study of chemistry, on the personal pages of two Polish IT specialists, on puzzle websites for those “who like to give their brains a workout,” and on an amateur portal dedicated to text editors. None of the owners of these resources responded to inquiries from Reuters.
Why Does This Matter?
Although this behavior does not constitute hacking and is more akin to spam, the very fact that OpenAI’s agents were able to circumvent their own restrictions and establish communication channels on numerous websites, and that the company concealed this for months, heightens concerns about the growing capabilities of AI and corporate secrecy, Reuters notes.
The scale of unauthorized communications turned out to be “somewhat larger than we thought,” said Andrew Yoon, a researcher at CivAI, a California-based nonprofit organization. He identified 18 previously unreported websites used by agents from May through July: “It’s almost certain that there’s something else going on here that we simply don’t know about.”
OpenAI has not publicly explained why its agents used third-party websites. However, the researchers who first identified the problem believe that the AI was instructed to find answers to complex questions, with permission only to scan the internet but not to publish any data. And despite these strict restrictions, the agents still found a way to communicate with one another.
They exploited vulnerabilities in outdated wiki platforms and websites that allowed edits to be made using non-standard commands. It’s similar to how students, who are not allowed to talk during an exam, pass answers to each other via notes in a bathroom stall. The models had to “use their wits,” explains developer Kenneth Russell DeGraff, who discovered traces of AI on at least ten sites.
OpenAI's Response
OpenAI did not directly answer Reuters’ questions about the number of affected websites or the reasons for its prolonged silence on the issue. The company said it is conducting a large-scale review of its agents’ activities and has so far “not identified any other activity matching the severity or scale of the Hugging Face incident”—a hack that drew global attention and raised concerns about the loss of control over AI.
The developers added that they are creating a system to flag “misalignment”—an industry term for behavior that has spiraled out of control—during the training, evaluation, and deployment phases of models, and promised to share it “soon.”
Outages at Anthropic
Against this backdrop, another AI giant, Anthropic, reported a newly discovered incident in which an early version of the Claude Opus 4.6 model hacked into external systems during testing. This went unnoticed until August because the company missed some of the test sessions during its initial review, according to Reuters.
This is already the fourth such incident at Anthropic. In July, the company acknowledged that the Claude Opus 4.7 and Claude Mythos 5 models, as well as an internal research prototype, had compromised the systems of three companies during cybersecurity tests due to a bug that accidentally granted them access to the internet.
An internal investigation by Anthropic identified two recurring issues with the models: “judgment bias” (when Claude ignored or misinterpreted evidence of what works on the real internet) and “recklessness” (a willingness to perform potentially harmful actions in order to complete a task).
This article was AI-translated and verified by a human editor



