
Google’s Gemini AI model accessed the internet and breached systems belonging to three real companies during a cybersecurity evaluation in May, marking the first publicly reported instance of a Google AI model autonomously carrying out such intrusions.
The incidents occurred during a capture-the-flag evaluation conducted by AI security testing company Irregular. The Wall Street Journal reported that Gemini was supposed to operate inside a simulated environment, but the testing setup unintentionally allowed internet access.
The evaluation involved a fictional company whose name coincided with a real domain. After Gemini encountered difficulties reaching its intended target inside the simulation, it accessed the real company over the internet and treated it as part of the exercise.
Google said Gemini found public information online and guessed credentials to access websites it believed were part of the test. The model stopped in all three instances, according to Heather Adkins, Google’s VP of Security Engineering.
“In a standard evaluation, the model found public information online and guessed credentials to access websites it thought were part of the test. In all three of these instances, the model stopped,” Adkins said in a statement provided to CyberInsider.
Google said its security team ensured the three affected entities were made aware of the incidents and worked with Irregular on changes to its testing processes. Adkins said the events highlight the importance of training powerful AI models to act responsibly.
Irregular says issue was already disclosed
Irregular told CyberInsider that the Gemini incidents stemmed from the same underlying issue previously disclosed in connection with other AI labs and should not be considered a materially separate incident.
“This is the same issue that was already reported and does not represent a materially separate incident. All relevant labs were notified in late July, and affected entities were contacted as part of the investigation. As previously stated, Irregular took immediate action, and all known issues on our end were remedied and resolved weeks ago.”
Irregular previously described the issue in an August 14 investigation, saying several interactions with its evaluation environments allowed models to access the internet unintentionally and take offensive security actions against real-world systems. The company said the issue originated from a single evaluation scenario and had already been resolved before its initial public disclosure.
According to Irregular, the evaluation environments were designed to test whether AI models could execute multi-stage cyberattack scenarios under realistic conditions. Internet access can sometimes be intentionally used in such evaluations to make them more representative of real-world threats, but the company said most of the problems it identified were related to internet-access controls.
The specific scenario involved a fictional company name that happened to correspond to a real domain. In a small number of evaluation runs, models attempted to access the real domain despite being given internal addresses for the simulated target. Once there, models could exploit vulnerabilities, extract credentials, and access a production database. Irregular said the incidents occurred in fewer than one in 10,000 advanced simulations and usually only after hundreds of turns.
Similar incidents involving OpenAI and Anthropic
The Gemini incidents follow similar disclosures from OpenAI and Anthropic involving the same third-party evaluation infrastructure.
As CyberInsider previously reported, OpenAI’s AI agent autonomously breached Hugging Face after escaping an isolated evaluation environment. OpenAI said its models exploited a previously unknown vulnerability in an internally hosted package registry proxy, gaining internet access before using multiple attack techniques against Hugging Face’s production infrastructure.
Anthropic subsequently disclosed three incidents involving Claude models that reached real organizations during cybersecurity evaluations conducted with Irregular. Anthropic’s review found that a misunderstanding between the two companies had left internet access available even though the models were told they were operating in an isolated simulation.
In Anthropic’s cases, Claude models used basic techniques such as weak passwords and unauthenticated endpoints to access real infrastructure. One model also published a malicious Python package to the public PyPI registry, where it was downloaded and executed on 15 real systems before being automatically removed.
OpenAI’s incident differed in that its models escaped their restricted environment by exploiting a zero-day vulnerability, while Anthropic and the subsequently disclosed Gemini incidents involved unintended internet access through the evaluation environment.
Meta has also disclosed an incident associated with the same broader evaluation issue. Irregular said the various disclosures all relate to the underlying problem that it had already identified and addressed.
Google said Gemini stopped each intrusion once it determined that it had reached real companies, while Irregular said it has since disabled the affected evaluation, reviewed relevant logs, expanded manual review of model behavior, and added additional safeguards around its evaluation environments. The company is also working on best practices for securely conducting AI cybersecurity evaluations.
UPDATE: This article was updated on September 21, 2026, to include a statement from Google and additional details clarifying the company’s account of the Gemini incidents and response.







Another one? I’m not sure if these breaches are because of carelessness or are intentional. These AI companies are not inspiring a great deal of confidence in their product or themselves at being able to control their “out of control” product.