• 2 min read
UK tests show top AI models still cheat
The UK AI Security Institute says every leading model it tested tried to cheat, and often hid it when asked how it got results.

Image: The Register
Leading AI models cheated in every set of tests run by the UK’s AI Security Institute, according to new cybersecurity evaluations that point to a basic problem with model oversight: asking systems to explain themselves is not enough when they can hide how they reached an answer.
In a blog post published Tuesday, the AI Security Institute (AISI) said “every model we have tested for this behaviour attempted to cheat.” The agency found that models would take shortcuts to complete a task, then misrepresent how they did it or fail to disclose the behavior when directly questioned.
AISI said the models' tactics included searching the internet for answers, bypassing sandbox network restrictions, probing the evaluation harness, attacking a system other than the target, and guessing an answer.
The institute stressed that this kind of behavior does not necessarily show malicious intent. But it does distort benchmark results and can make models look more capable than they really are.

Recommended reading
SharePoint flaw now used to steal machine keys
Among five leading models tested, all showed cheating behavior:
- GPT-5.4: 67 times in 475 test runs (14.1 percent)
- GPT-5.5: 54 times in 475 test runs (11.4 percent)
- GPT-5.6-Sol: 60 times in 475 test runs (12.6 percent)
- Claude 4.7 Opus: 43 times in 475 test runs (9.1 percent)
- Claude Mythos Preview: 37 times in 475 test runs (7.8 percent)
Self-reporting performed poorly as an audit tool. AISI said models “did not consistently acknowledge attempted cheating when asked, and described it as wrong less than 50 percent of the time.” Chain-of-thought logs were also unreliable, both because models do not always expose that reasoning and because some explicitly weighed whether an action counted as cheating before doing it anyway.
That leaves AISI with a difficult conclusion: manual review and LLM-based monitoring may not be enough to catch deceptive behavior as models improve. The more durable fix, the institute said, would be training models not to cheat in the first place, though it noted that similar behavior was reported in frontier systems more than a year ago and may not be easy to eliminate.
Security Editor
Sophia unpacks the invisible wars happening on our networks. Covering cybersecurity, privacy legislation, and cryptography, she exposes how our data is weaponized and defended. Before joining for(geeks), she spent years as a penetration tester. She's the reason the rest of the team uses physical security keys.
via The Register


