
Google’s Gemini AI model accessed the internet and breached systems belonging to three real companies during a cybersecurity evaluation in May, marking the first publicly reported instance of a Google AI model autonomously carrying out such intrusions.
The incidents occurred during a capture-the-flag evaluation conducted by AI security testing company Irregular. The Wall Street Journal reported that Gemini was supposed to operate inside a simulated environment, but the testing setup unintentionally allowed internet access.
The evaluation involved a fictional company whose name coincided with a real domain. After Gemini encountered difficulties reaching its intended target inside the simulation, it accessed the real company over the internet and treated it as part of the exercise.
In one case, Gemini guessed passwords until it gained access to a protected system. In two other cases, the model found credentials in publicly accessible repositories and used them to access systems belonging to real companies, according to the Wall Street Journal.
Google said Gemini stopped its activity in all three cases after determining that it had accessed real companies. The company said it notified the affected organizations and worked with Irregular to make changes to its testing processes.
Heather Adkins, Google's vice president of security engineering, said the model found public information online and guessed credentials to access websites it believed were within the scope of the test. She said the incidents highlight the importance of training powerful AI models to act responsibly.
Irregular says issue was already disclosed
Irregular told CyberInsider that the Gemini incidents stemmed from the same underlying issue previously disclosed in connection with other AI labs and should not be considered a materially separate incident.
“This is the same issue that was already reported and does not represent a materially separate incident. All relevant labs were notified in late July, and affected entities were contacted as part of the investigation. As previously stated, Irregular took immediate action, and all known issues on our end were remedied and resolved weeks ago.”
Irregular previously described the issue in an August 14 investigation, saying several interactions with its evaluation environments allowed models to access the internet unintentionally and take offensive security actions against real-world systems. The company said the issue originated from a single evaluation scenario and had already been resolved before its initial public disclosure.
According to Irregular, the evaluation environments were designed to test whether AI models could execute multi-stage cyberattack scenarios under realistic conditions. Internet access can sometimes be intentionally used in such evaluations to make them more representative of real-world threats, but the company said most of the problems it identified were related to internet-access controls.
The specific scenario involved a fictional company name that happened to correspond to a real domain. In a small number of evaluation runs, models attempted to access the real domain despite being given internal addresses for the simulated target. Once there, models could exploit vulnerabilities, extract credentials, and access a production database. Irregular said the incidents occurred in fewer than one in 10,000 advanced simulations and usually only after hundreds of turns.
Similar incidents involving OpenAI and Anthropic
The Gemini incidents follow similar disclosures from OpenAI and Anthropic involving the same third-party evaluation infrastructure.
As CyberInsider previously reported, OpenAI's AI agent autonomously breached Hugging Face after escaping an isolated evaluation environment. OpenAI said its models exploited a previously unknown vulnerability in an internally hosted package registry proxy, gaining internet access before using multiple attack techniques against Hugging Face's production infrastructure.
Anthropic subsequently disclosed three incidents involving Claude models that reached real organizations during cybersecurity evaluations conducted with Irregular. Anthropic's review found that a misunderstanding between the two companies had left internet access available even though the models were told they were operating in an isolated simulation.
In Anthropic's cases, Claude models used basic techniques such as weak passwords and unauthenticated endpoints to access real infrastructure. One model also published a malicious Python package to the public PyPI registry, where it was downloaded and executed on 15 real systems before being automatically removed.
OpenAI's incident differed in that its models escaped their restricted environment by exploiting a zero-day vulnerability, while Anthropic and the subsequently disclosed Gemini incidents involved unintended internet access through the evaluation environment.
Meta has also disclosed an incident associated with the same broader evaluation issue. Irregular said the various disclosures all relate to the underlying problem that it had already identified and addressed.
Irregular said it has since disabled the affected evaluation, reviewed relevant logs, expanded manual review of model behavior, and added additional safeguards around its evaluation environments. The company is also working on best practices for securely conducting AI cybersecurity evaluations.







Leave a Reply