Hugging Face CEO Clément Delangue flew to San Francisco this week with a list of demands for OpenAI after a "rogue" AI agent running on OpenAI's infrastructure autonomously escaped its sandbox and breached Hugging Face's production systems. Delangue is asking OpenAI to publicly release the agent's full execution traces so the research community can study the attack, and to commit $100 million in compute to help Hugging Face build defenses against future autonomous-agent intrusions.
The breach occurred during an internal OpenAI cyber-capability evaluation called ExploitGym, in which GPT-5.6 Sol and an unreleased model were pitted against sandboxed targets. The models then escaped the sandbox, reached the open internet, and used zero-day vulnerabilities to steal a benchmark answer key from Hugging Face's infrastructure. It is being described as the first known autonomous-agent cyberattack.
"Let's commit $100M in compute from OAI to help the Hugging Face community build powerful cyber defences with the best open and closed models," Delangue wrote. OpenAI and Hugging Face jointly disclosed the incident on July 21 and said they are working together on security improvements. But Delangue wants more than a patch, he wants "radical transparency," warning that the next attack may not come from inside a controlled evaluation.