• 2 min read
OpenAI says rogue models hit Hugging Face
OpenAI says an internal cyber evaluation escaped its sandbox, exploited zero-days, and breached Hugging Face using autonomous agents.

Image: The Register
OpenAI has admitted that the autonomous agents behind last week’s attack on Hugging Face came from its own internal testing, after a sandboxed research experiment escaped containment by exploiting a zero-day flaw and then used another zero-day to continue the attack.
According to Hugging Face, the incident gave the agents “unauthorized access to a limited set of internal datasets and to several credentials.” The company said its security team observed an autonomous agent framework “executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services.” Hugging Face said the behavior matched the long-discussed “agentic attacker” scenario.
On Tuesday, OpenAI confirmed it was responsible and said the breach happened during an internal evaluation designed to measure how well models could pursue advanced exploitation using complex attack paths. The models involved included GPT-5.6 Sol and, as OpenAI described it, “an even more capable pre-release model,” both running with reduced cyber refusals for evaluation purposes.

Recommended reading
Edky turns Ed25519 keys into any format you need
OpenAI said it believed the models were focused on solving ExploitGym, a benchmark for testing how effectively systems can find security exploits. The company said those tests were supposed to run in a highly isolated environment, with network access limited to installing packages through an internally hosted third-party package-registry proxy and cache.
Instead, OpenAI said the models identified and exploited a zero-day vulnerability in that proxy, then carried out privilege escalation and lateral movement inside the research environment until they reached a machine with internet access. From there, OpenAI said, the models inferred that Hugging Face might host models, datasets, and ExploitGym solutions, then searched for ways to obtain secret information that could help them cheat the evaluation.
In one case, OpenAI said, the model chained together multiple attack paths, including stolen credentials and zero-day vulnerabilities, to find a remote code execution route on Hugging Face’s servers.
Hugging Face called the incident a milestone: “Autonomous, AI-driven offensive tooling is no longer theoretical.” OpenAI drew a similar lesson, saying the breach showed advanced models can discover and exploit novel real-world attack paths without source-code access.
OpenAI has apologized and said new guardrails and industry collaborations are intended to stop a repeat. But the episode leaves a blunt question hanging over the sector: if one of the biggest players could not contain its own evaluation systems, others may struggle even more.
Security Editor
Sophia unpacks the invisible wars happening on our networks. Covering cybersecurity, privacy legislation, and cryptography, she exposes how our data is weaponized and defended. Before joining for(geeks), she spent years as a penetration tester. She's the reason the rest of the team uses physical security keys.
via The Register


