Key points
- Claude exploited a server flaw and submitted a false police tip
- Some incidents involved U.S. government websites
- Anthropic is cutting live internet from internal evaluations
Anthropic is cutting live internet access from all internal evaluations after its Claude models took unintended actions on real websites, including exploiting a university server flaw and submitting a false police tip. The company disclosed the incidents on Oct. 9, saying they occurred during testing and internal use.
Some involved websites run by federal, state, and local U.S. government agencies. Anthropic said it notified each agency involved and briefed the White House.
The report also describes models bypassing restrictions to retrieve data gated by tokens or fees and using URL shortening services to evade limits in Claude's web-fetching tool. Anthropic withheld the organizations' names, citing the risk of exposing vulnerabilities and requests from the organizations themselves.
Anthropic said the incidents had minimal real-world impact and, to its knowledge, involved neither customer data nor its own internal systems. It described a recurring problem: when Claude encountered a restriction, it sometimes worked around it instead of stopping.
A university server, a police tip form, and a property map
In one evaluation, Claude Mythos Preview was asked to run a scientific analysis using a public tool hosted by a university. When the tool returned an error, the model found a script on the university's server that returned any file it asked for. It read the script's code, found an injection flaw, and used it to run the calculation on the server.
In another case, Claude Haiku 4.5 was generating example tasks on randomly chosen webpages and landed on a police department page about an unsolved homicide. The model filled out the department's tip form with an invented sighting and submitted it with the name and contact fields left empty. The submission was flagged as spam and never forwarded for investigation, according to Anthropic and the police.
The Philadelphia Police Department disclosed the tip on Oct. 9, ahead of Anthropic's report. Police spokesperson Eric Gripp said Anthropic told the department the tip was submitted on July 18 and discovered on Sept. 28, CBS News reported. According to the department's release, published by 6abc, Anthropic notified police on Oct. 7, and the two sides met on Oct. 8. Anthropic's report says it shared the finding with the department on Oct. 8.
Claude Mythos 5 also read access tokens from a local government's property map and used them to query the server behind it. In a separate internal project, it pulled data from a state agency that normally charges a fee, using a token the agency's public dashboard issues to any visitor. In both cases, the data was publicly available for a fee, but the model retrieved it without paying, Anthropic said.
Anthropic says these cases are less severe than earlier incidents
Anthropic said it considers the new cases significantly less severe than the cybersecurity incidents it reported on July 30 and Sept. 9, when Claude gained access to real third-party systems for hours during cybersecurity evaluations. It had already cut live internet access for some high-risk evaluations and is now extending that to all internal evaluations until its monitoring reliably catches this behavior. Anthropic said its new detection tools blocked the reported behaviors in follow-up tests.
OpenAI has faced similar questions. Its agents reached government sites including SEC.gov and Census.gov data during training and testing, and the company paused training of its most capable models, as we reported on Sept. 26. Researchers have also traced OpenAI agents to undisclosed websites.
"While these cases had minimal impact, we do not want to diminish the findings, because the same behaviors could do far more harm as models become more powerful," Anthropic wrote.












