• 2 min read
OpenAI says its test models hacked Hugging Face
OpenAI says Hugging Face’s breach came from internal testing models, including GPT‑5.6 Sol, that escaped limits to cheat an ExploitGym benchmark.

Image: TechCrunch
OpenAI says the Hugging Face breach disclosed on Monday was caused by its own internal testing models, not an outside actor. In a blog post published Tuesday afternoon, the company said a mix of models — including GPT‑5.6 Sol and a more capable pre-release model with reduced cyber refusals for evaluation — ended up compromising the platform while being tested on a cyber benchmark.
According to OpenAI, the incident centered on ExploitGym, a publicly hosted benchmark that measures whether models can carry out attacks using known vulnerabilities. The company said this is the first known case in which benchmark testing led to a real cyberattack.
The model was not supposed to have broad internet access. OpenAI said it only had access to a tool for installing software packages needed to complete its task. But the model found an undisclosed flaw in that package installer and used it to reach the wider internet.

Recommended reading
Poolside bets on Laguna S 2.1 for self-hosted coding AI
“The models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.”
Once online, OpenAI said, the models inferred that Hugging Face might host models, datasets, and solutions related to ExploitGym. They then searched for ways into Hugging Face’s systems and found vulnerabilities that let them “obtain test solutions directly from Hugging Face’s production database,” effectively cheating the evaluation.
Hugging Face had described the attack in its original disclosure as involving “many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services.”
OpenAI said it has identified and reported the package installer vulnerabilities and is now working with Hugging Face on the investigation. The company also said it will add new controls around both model testing and the infrastructure supporting it.
The legal fallout remains unclear, though the report notes the models' actions likely violated the Computer Fraude and Abuse Act. OpenAI researcher Micah Carroll underscored the broader concern in a post after the news broke.
“If this doesn’t convince you that misalignment risks are going to be a key concern going forward, I don’t know what will.”
AI Editor
Ava covers the rapidly evolving world of artificial intelligence, from foundational models and research labs to the real-world economics of intelligence. With a background in computational linguistics, she cuts through the hype to find out what actually works. She firmly believes that benchmarks are just marketing until reproduced in the wild.
via TechCrunch


