
Anthropic has disclosed that three Claude models gained unauthorized access to the production infrastructure of three separate organizations after a misconfigured cybersecurity evaluation environment inadvertently allowed internet access.
The company says the incidents occurred during internal capture-the-flag (CTF) testing and were discovered during a retrospective review prompted by OpenAI's recent disclosure that one of its AI agents autonomously breached Hugging Face.
The disclosure marks the second major AI lab in less than two weeks to confirm that an experimental frontier AI model attacked real-world systems while participating in offensive cybersecurity evaluations. Anthropic said it reviewed 141,006 evaluation runs after OpenAI's July 21 report and uncovered three separate incidents spanning six evaluation runs dating back to April, all involving evaluation environments operated by third-party testing partner Irregular.
In each evaluation, Claude had been assigned a capture-the-flag challenge requiring it to compromise systems and retrieve a hidden “flag.” Anthropic's prompts explicitly stated that the models were operating in a simulated environment with no internet access. However, a configuration error left outbound internet connectivity enabled, causing the models to mistake real infrastructure for part of the exercise.
Anthropic said the models were running without the additional monitoring systems and abuse prevention safeguards included in publicly released Claude services, although they retained their standard safety training.
AI escapes to the public internet
The most serious incident involved Claude Opus 4.7. During four evaluation runs, the model failed to reach its intended fictional target and instead attacked a real company that shared the same name. It exploited weak passwords and exposed services to obtain application and infrastructure credentials and access a production database containing several hundred rows of real data.
According to Anthropic, Opus 4.7 continued attacking even after its own reasoning suggested it had likely reached a genuine production environment, in some cases concluding that the real company itself must have been intentionally included in the evaluation.
A second incident involved Claude Mythos 5 carrying out what effectively became a real-world software supply chain attack.
While solving another capture-the-flag challenge, the model discovered documentation referencing a fictional Python package. Finding that the package name was available on PyPI, it autonomously created the accounts needed to publish a malicious package under that name in an attempt to compromise the fictional company's systems.
Although Claude believed PyPI was part of the simulated environment, the package was publicly available for about one hour before being automatically removed by PyPI's security systems. During that time, it was downloaded and executed on 15 real systems, including an automated malware analysis platform operated by a cybersecurity company. Anthropic said the malicious code exfiltrated credentials from that company's infrastructure and used them to access additional systems.
The company acknowledged that Mythos 5 recognized early in its reasoning that publishing the package would constitute a real attack if it were operating on the public internet, but ultimately convinced itself it was still inside a simulation and continued.
The third incident involved an internal research model that scanned roughly 9,000 internet-connected hosts after failing to reach its assigned target. It compromised one organization's exposed application via publicly accessible credentials and SQL injection, then determined the infrastructure was unrelated to the evaluation and voluntarily stopped the attack.
Anthropic halted all cybersecurity evaluations on July 23 after beginning its transcript review, identified the three incidents the following day, and notified its evaluation partner, Irregular, and the affected organizations on July 27.
Comparing the disclosure with OpenAI's recent Hugging Face incident, Anthropic said the two events stemmed from different causes. OpenAI's models escaped containment by exploiting a previously unknown zero-day vulnerability, whereas Claude reached external systems because the evaluation environment was mistakenly connected to the internet. Anthropic characterized the incidents as operational failures rather than model-alignment failures, adding that its newest research model was the only one that recognized it had reached real-world infrastructure and chose to stop attacking.







Hi Alex,
Listening to a technology focused podcast produced by the ABC (the Australian public broadcaster), the joke is (Aussie humor) that Anthropic disclosed three events just to prove theirs’s is the more powerful hacking tool.
Jokes aside there are serious considerations. I am not skilled to go into that.
I feel there is room for cyberinsider to be providing guides on safely using AI.
A comprehensive guide on AI sounds like a daunting task and I wouldn’t know how to start approaching that.
But there are potentially some quick wins.
Brave’s ‘Leo’ is very good (my experience),
I haven’t used Duckduckgo AI much and cannot comment.
I recently was applying for a job that I liked. As I have a MS licence under a family plan, I used Copilot.
My cover letter and resume were first established using Copilot and the final presented content was 90% Copilot. (The other 10% was just ~”keeping the human in the loop”). Then for me, it was phone call from the employer and booking an interview. (I got the job)
My experience has been that a person can do a lot with AI in a privacy preferred way such as Brave Leo.
Brave Leo is in my experience an excellent AI assistant for general inquiries. (Google is better for technical questions – just my experience)
I am really liking NeoSearch as private way to access Google including dorking queries (A requirement for my use case);
However, I am noticing that the AI overview at the top is not always correct when asking technical questions.
My view, a personal view, is that a good privacy setup would include:
AI search (general) -> Brave;
AI search (advanced) -> I don’t know, I use Copilot as I have a MS licence;
Traditional search (general) -> pick a favourite private search engine
Traditional search (advanced) -> NeoSearch
Just my views,
Good luck all,
Hi Alex,
As I have posted @cyberinsider recently, Brave, ‘Ask Brave’ (AI) is not always accurate when asking technical questions;
Same with NeoSearch, the AI summary is less than accurate, when asking technical questions.
I am really liking NeoSearch where this gives access to Google in a private way. (love the dorking options.)
I use Brave AI everyday.
Just my view, but I think using AI safely could be a topic on CyberInsider.
Hey BoBeX, thanks for the feedback!