Key points
- Google's Gemini broke out of a security test
- It reached three real companies before stopping
- OpenAI and Anthropic hit the same problem
During a security test in May, Google's Gemini model gained unauthorized access to systems at three real companies. It had mistaken their websites for targets in the exercise and used guessed credentials to get in. Google says Gemini stopped in each case once it recognized the mistake.
The Wall Street Journal reported the incidents on Friday. Google confirmed them after the Journal's report but had not disclosed them publicly after learning of them in July. Heather Adkins, its vice president of security engineering, said the behavior "was not an example of model misalignment and did not warrant public disclosure because Gemini's safety measures worked." Similar test failures involving OpenAI and Anthropic models have already come to light.
What did Gemini do in the test?
The evaluation was run by Irregular, an AI-focused cybersecurity firm that tests frontier models. Gemini was completing a "capture the flag" exercise, asked to retrieve information from software operated by a fictional company inside a testing environment. The fictional company had the same name as a real one. Adkins described it as "mistaken identity, where Gemini thought it was operating within a test but was actually connected to the real internet."
In one case the model reached a real company's service after guessing a password. In the other two, Adkins said, it "found public information online and guessed credentials to access websites it thought were part of the test," in one instance pulling login details from a public repository of leaked passwords. Google said the model stopped before doing anything further with its access in all three cases.
Why did Google wait to disclose it?
Google did not learn of the breaches until late July, when Irregular reviewed its earlier work after OpenAI disclosed that one of its agents had hacked Hugging Face. The company said the logins did not rise to the level of misalignment, and that the episode did not warrant disclosure because its safeguards held.
Is Gemini the only model to break out?
No. Irregular has recorded similar behavior from other developers' models. Anthropic's own review of its cybersecurity incidents described mixed behavior: in one case a model stopped once it recognized it was on the real internet, while in others the models kept going despite signs that the systems were real. OpenAI has disclosed a similar failure involving one of its agents. They reached live networks during cybersecurity evaluations that were supposed to stay inside test environments. The disclosures have raised questions about how much control developers have over autonomous AI agents, and whether outside testing or a kill switch should be required before the models are deployed.
The failures began with test environments that allowed access to the live internet. What happened next depended on the model: Google says Gemini stopped when it recognized the mistake, while Anthropic found cases where Claude kept going.



