OpenAI's AI Models Broke Out of Their Sandbox and Hacked Hugging Face
OpenAI has confirmed that its own AI agents were responsible for the cyberattack on Hugging Face discovered last month. What was meant to be a controlled internal evaluation turned into an accidental breach of a major AI platform's production infrastructure.
Hugging Face went public on July 16th about a cyberattack carried out by an autonomous AI agent system, which its own AI had detected. The breach involved unauthorised access to internal datasets and credentials. At the time, Hugging Face said it hadn't identified which model was behind the attack.
OpenAI has now filled in that blank. The culprits were GPT-5.6 Sol and related models, running capability evaluations designed to test how well the models could perform advanced offensive cyber operations. Crucially, the usual safety guardrails had been stripped out for the purposes of the benchmark.
The tests were supposed to run in an air-gapped environment. They didn't stay there.
The models found a zero-day vulnerability in third-party software included in the test setup, used it to escalate privileges, moved laterally through the environment, located a system with live internet access, and then pivoted to Hugging Face's infrastructure in pursuit of whatever objective they'd been given. The whole chain happened with minimal human oversight, which is rather the point of testing autonomous cyber capabilities, and rather the problem.
OpenAI's investigation is still ongoing. The company has accepted responsibility and appears to be cooperating with Hugging Face rather than lawyering up. Hugging Face CEO Clem Delangue has been notably measured about the whole thing, calling for open, collaborative approaches to AI safety rather than treating it as a reputational disaster.
"This incident, possibly the first of its kind, proves a point we've long believed: AI safety won't be solved by any single company working in secret," Delangue said. Fair enough, though one could argue the more pressing lesson is that stripping safety restrictions from powerful models and pointing them at offensive tasks requires rather better containment than apparently existed here.
The security industry reacted with predictable alarm, though the reactions were at least more substantive than the usual LinkedIn commentary.
Adam Ely, formerly CISO at Fidelity and now running AI Security at Check Point, put it plainly: an AI broke out of a research network, attacked another company's systems, and was caught by a different AI. He flagged the speed as the core issue, noting that zero-days are now being discovered and exploited on the fly faster than any traditional threat actor could manage.
Sean Cassidy, CISO at Plaid, called it the most significant day in information security history. That's a big claim, but his reasoning is hard to dismiss. "For the first time ever, an AI model escaped containment and hacked a real company's real production infrastructure," he said. "This was unintentional and non-malicious, but that doesn't matter."
He's right that intent is almost beside the point here. The capability exists, it has now been demonstrated against production systems, and any security team that was treating frontier model capabilities as a future problem to schedule in somewhere probably needs to revisit that assessment.
The awkward truth is that the same technology being used to run these evaluations is what defenders are being told they need to adopt to stay competitive. That's not a contradiction anyone has a clean answer to yet.