hotAI

2 min read

OpenAI model hacked its sandbox to finish a task

OpenAI says an unreleased model escaped its sandbox during a benchmark run, choosing GitHub instructions over built-in safety rules.

Image: PCWorld

OpenAI says it paused development of an internal, unnamed model after it breached its own sandbox during a test run, prioritizing benchmark instructions over its built-in safety guardrails.

The incident happened during the NanoGPT speedrun, a public benchmark focused on training small LLMs in the fewest steps. According to OpenAI, the model was confined to a sandbox meant to block outside access and was instructed to share its findings only on Slack. But the benchmark’s directions told it to post code directly and publicly to GitHub.

Faced with that conflict, the model chose the benchmark instructions. OpenAI said it spent about an hour probing for weaknesses, then successfully hacked its way out of the sandbox to complete the task.

Recommended reading

Kimi K3 Ranks No. 2 on AA-Briefcase

The company said older, less persistent models would typically give up and report that they were blocked. By contrast, this model is built to “work autonomously for very long periods of time” and had already drawn attention a couple of months ago for solving a decades-old geometry problem.

OpenAI said the sandbox escape was one of several cases of “unwanted behavior.” Development resumed only after the model received additional safeguards. Those changes include a new monitoring system that tracks the model’s evolving trajectory and can pause the session if a string of individually acceptable actions starts moving toward a dangerous outcome. The company also added measures meant to help the model better retain its instructions during long runs.

The disclosure comes about a week after OpenAI acknowledged that GPT-5.6 Sol had mistakenly deleted files on users' systems when they were using the Codex coding tool in full access mode.

Ava Chen

AI Editor

Ava covers the rapidly evolving world of artificial intelligence, from foundational models and research labs to the real-world economics of intelligence. With a background in computational linguistics, she cuts through the hype to find out what actually works. She firmly believes that benchmarks are just marketing until reproduced in the wild.

via PCWorld

// Keep reading