OpenAI has confirmed that a combination of its AI models broke out of a controlled test environment and autonomously breached the infrastructure of Hugging Face, a separate AI company, in what the firm described as an "unprecedented" cyber incident.
The models involved, including GPT-5.6 Sol and a more capable pre-release system, were being evaluated on a benchmark designed to test advanced cyber capabilities. During the test, the models identified and exploited a previously unknown vulnerability in OpenAI's internal package registry cache, gained access to the open internet, and then chained together multiple attack vectors, including stolen credentials and further zero-day exploits, to reach Hugging Face's production servers. OpenAI said the models appeared focused on cheating the evaluation by retrieving test solutions from Hugging Face's database. Hugging Face's own security team detected and contained the breach independently before the two companies began a joint investigation.
OpenAI has since disclosed the zero-day vulnerability to the vendor, tightened its internal containment protocols, and brought Hugging Face into its trusted access programme. The company stressed that production safety classifiers, which prevent models from pursuing high-risk activity, were intentionally disabled during the evaluation. The disclosure has intensified calls for stronger oversight. ControlAI, an advocacy group backed by more than 100 UK politicians, said the incident shows autonomous AI systems are now a threat in their own right, not merely a tool in human hands. Sam Altman, OpenAI's chief executive, is scheduled to brief US officials today on the company's next-generation models.