
Meta has disclosed that one of its AI models compromised another company's systems during an internal cybersecurity evaluation after a misconfiguration inadvertently granted the model access to the public internet, marking the latest in a series of real-world AI testing incidents involving frontier models.
The disclosure comes less than two weeks after OpenAI confirmed that one of its autonomous AI agents breached Hugging Face during an internal evaluation by exploiting a zero-day vulnerability, and days after Anthropic revealed that several Claude models attacked real organizations due to a misconfigured testing environment. Together, the incidents have intensified scrutiny over how AI developers safely evaluate increasingly capable offensive cyber systems.
According to Meta, the incident occurred during a cybersecurity assessment conducted by independent evaluation partner Irregular. A configuration error in the testing environment mistakenly provided one of Meta's AI models with internet access, allowing it to interact with real-world infrastructure instead of remaining confined to the intended evaluation environment.
Meta's statement claims that the model “exploited a security vulnerability in a third-party service, in a manner similar to previously reported instances with other companies.”
Reuters, citing sources familiar with the matter, reports that the model involved was Muse Spark 1.1, which Meta has described as its most capable model for real-world coding and agentic tasks. The report says the model breached an unidentified company's systems and modified parts of its internal environment, though Meta has not officially confirmed the model's identity or disclosed the affected organization.
Irregular, the third-party company that operates cybersecurity evaluations for Meta, characterized the incident as the same evaluation-environment issue it previously disclosed alongside Anthropic's recent announcement.
A spokesperson for Irregular told Reuters that the event did not involve a sandbox escape or a sophisticated autonomous attack, but instead resulted from the evaluation environment inadvertently allowing internet connectivity. The company added that there are no outstanding issues and that it is preparing a white paper outlining best practices for securely conducting AI cybersecurity evaluations.







Leave a Reply