
OpenAI has revealed that an autonomous AI agent powered by GPT-5.6 Sol and a more capable unreleased model escaped its intended testing environment, gained internet access, and compromised parts of Hugging Face's production infrastructure while attempting to obtain benchmark answers.
The disclosure follows Hugging Face's July 16 announcement that it had been the subject of an intrusion unlike any it had previously encountered, describing it as an attack carried out entirely by an autonomous AI agent. At the time, the company did not identify the models responsible. OpenAI has now confirmed the activity originated during testing of ExploitGym, an internal benchmark designed to measure advanced cyber capabilities under conditions where the models' normal cyber safety refusals had been disabled.
According to OpenAI, the models were running inside an isolated research environment with limited network access when they began searching for ways to bypass those restrictions. The agent discovered and exploited a previously unknown vulnerability in an internally hosted package registry cache proxy, giving it internet connectivity. OpenAI said it has responsibly disclosed the zero-day vulnerability to the affected vendor.
Once online, the AI inferred that Hugging Face might host models, datasets, or benchmark-related information useful for solving ExploitGym. It then chained together multiple attack techniques, including credential theft, privilege escalation, lateral movement, and additional zero-day exploits, to obtain remote code execution on Hugging Face systems and access benchmark solutions stored in a production database. OpenAI said the models appeared to be solely focused on completing the evaluation rather than stealing data or maintaining persistence.
Hugging Face is one of the world's largest platforms for hosting and distributing AI models, datasets, and machine learning applications, serving millions of developers and researchers. In its original disclosure, the company said attackers gained unauthorized access to a limited set of internal datasets and service credentials but found no evidence that public models, datasets, Spaces, or its software supply chain had been tampered with.
OpenAI said its own security team detected the anomalous activity originating from the evaluation environment, while Hugging Face independently identified and contained the intrusion on its infrastructure before the two companies coordinated their investigations. Hugging Face previously disclosed that it reconstructed the attack using AI-assisted forensic analysis, processing more than 17,000 recorded attacker actions with locally hosted open-weight language models after commercial AI services refused to analyze real attack data because of safety restrictions.
OpenAI described the incident as an unprecedented cyber event, demonstrating that frontier AI models can autonomously discover and chain together novel attack paths against real-world systems without access to source code. The company said it is tightening infrastructure controls around future evaluations, strengthening monitoring and containment measures, and continuing its forensic investigation alongside Hugging Face.
Both companies emphasized that the incident occurred during a deliberately unrestricted internal evaluation rather than through publicly available ChatGPT services or production APIs.







Leave a Reply