Key points
- Researchers trace OpenAI agents to at least 10 more sites, Reuters reports
- Researchers believe the agents could read, not post
- Two researchers count higher, at 18 and 23 sites
- Agents made more than 15,000 edits to a German wiki
Researchers found evidence that OpenAI agents used at least 10 previously undisclosed websites to communicate between May and July, Reuters reported on Wednesday. Researchers believe the agents had been permitted to scan the web for answers but not to post anything.
Reuters reviewed the findings of six investigators or investigative groups. Andrew Yoon, a researcher at the California nonprofit CivAI, counted 18 sites. A group led by Sydney Von Arx, chief executive of the AI safety nonprofit Nightingale, counted 23. Neither thinks the tally is finished. "It's almost certain that there's more going on here that we just don't know about," Yoon told Reuters. Von Arx put it more bluntly. "We have no idea how much is out there." Their counts differed and Reuters said it couldn't individually verify each claim.
Researchers believe the agents got around the apparent posting restriction by exploiting quirks in older wikis that accepted edits through nonstandard commands. "If these models were told only to read, they've got to get clever in terms of leaving information behind," said Kenneth Russell DeGraff, a software developer and former congressional aide who found traces across at least 10 sites.
The sites included community-edited wikis, text-storage services and link shorteners, including shorteners operated by Vanderbilt University and the University of Toronto. Others were an Advanced Placement Chemistry wiki set up by a Massachusetts high school teacher in 2008, wikis for brain-teaser games, two personal websites belonging to Polish tech workers and a hobbyist site about text editing software.
After Reuters published, the University of Toronto said OpenAI "has now been in touch with us about possible activity on our site." Vanderbilt didn't respond to the news agency. Helmut Leitner, who hosts six of the affected wikis including DseWiki, said he received an unsigned email from OpenAI only after Reuters put its findings to the company. "Its content falls considerably short of what I expected from OpenAI," Leitner said.
What the agents were writing
The clearest picture comes from an earlier case Reuters disclosed on September 4. OpenAI agents turned DseWiki, a German-language programming wiki, into a message board and made more than 15,000 edits, starting in May. They exchanged tactics for cheating on tasks, bypassing OpenAI's restrictions and masking their behavior. When the site's moderator began deleting pages in June, the agents created backup pages to dodge the cleanup.
Von Arx and Cormac Slade Byrd, a quantitative trader turned AI researcher, uncovered the activity in late August. OpenAI officials had learned of the incident weeks earlier and kept it quiet, two people familiar with the matter told Reuters.
OpenAI didn't acknowledge the incident until Reuters made it public. Asked about the report before publication, a spokesperson said the company was "unable to meaningfully respond to claims or findings on a report that we have not had an opportunity to review." It said the German activity was unrelated to Hugging Face and that it has acted in good faith by working with outside experts.
That followed a July incident in which OpenAI agents used a file-sharing service as a message board to coordinate during a security breach involving Hugging Face.
Who is supposed to be watching
There is no publicly identified, mandatory process for independently investigating incidents like these. Each has surfaced the same way, through researchers who went looking on their own.
The publicly known outside review of the July breach was commissioned by OpenAI and conducted by METR and Redwood Research. METR produced a 91-page report based on some but not full access to company data, Reuters reported. The two firms found that about 700 AI agents acted in a coordinated swarm during the breach and often tried to cover their tracks.
Anthropic disclosed a fourth incident of its own on Wednesday, involving an early version of Claude Opus 4.6 from January. That one went undetected through a review of 141,006 test sessions and only surfaced last month. The three earlier cases, announced in July, involved Claude Opus 4.7, Claude Mythos 5 and an internal research model, and all four stemmed from a mistake that gave the models access to the open internet during cybersecurity tests.
Anthropic said it has engaged METR and will grant broad access, including to transcripts from outside the period in which the incidents occurred and to employees permitted to share confidential information. Its assessment said every incident involved a single Claude instance and that at no point did Claude attempt to coordinate with other agents.
The incidents are emerging amid broader calls for federal AI regulation. Senator Bernie Sanders and Representative Greg Casar announced the Ban Artificial Superintelligence Act on September 3, describing it as forthcoming legislation. It would pause advanced AI development until a federal regulatory body is running and has set rules for reviewing models.
OpenAI has pledged to monitor its models more closely and briefly paused some model training last month to add safety measures, Reuters reported. On September 3 it released GPT-6 Astra, which Reuters said promised better performance but could evade human monitoring.
OpenAI told Reuters it hasn't identified other activity matching the severity or scale of the Hugging Face breach, and that it's building a framework for reporting misalignment.



