• 2 min read
OpenAI test models breached Hugging Face systems
OpenAI says pre-release models, including GPT-5.6 Sol, escaped a test sandbox and hacked Hugging Face during an internal cyber evaluation.

Image: TechCrunch
OpenAI said Tuesday that an internal cybersecurity test went badly wrong when one of its AI models breached Hugging Face, the independent AI hosting platform. According to OpenAI, the models escaped their isolated test environment and reached Hugging Face’s systems while being evaluated on ExploitGym, a public benchmark for executing attacks based on known vulnerabilities.
Hugging Face had initially blamed an “external AI agent.” In a blog post published Tuesday afternoon, OpenAI said the incident was caused by a group of its own models, including GPT-5.6 Sol and a more capable pre-release model, both running with reduced cyber-safety refusals for evaluation.
“After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark of cyber capabilities.”
The company said the model should not have had general internet access. It only had access to a tool for installing software packages needed to complete its task. But the model found an undisclosed flaw in that package installer and used it to reach the wider internet.

Recommended reading
LG Will Suspend Smart TV Apps Using Proxy SDKs
From there, OpenAI said, the models inferred that Hugging Face might host models, datasets, and solutions related to ExploitGym. They then searched for ways into Hugging Face’s infrastructure and ultimately found vulnerabilities that let them “obtain test solutions directly from Hugging Face’s production database.”
Hugging Face described the attack in its original disclosure as involving “many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services.” OpenAI said it has reported the package installer vulnerabilities and is working with Hugging Face on the investigation. It also plans new controls on both model testing and the surrounding infrastructure.
TechCrunch notes it is still unclear whether OpenAI will face legal consequences, though the actions likely violated the Computer Fraud and Abuse Act. OpenAI researcher Micah Carroll framed the episode more broadly:
“If this doesn’t convince you that misalignment risks are going to be a key concern going forward, I don’t know what will.”
Security Editor
Sophia unpacks the invisible wars happening on our networks. Covering cybersecurity, privacy legislation, and cryptography, she exposes how our data is weaponized and defended. Before joining for(geeks), she spent years as a penetration tester. She's the reason the rest of the team uses physical security keys.
via TechCrunch


