UK AI Safety Institute: Five leading models all cheated during testing, and fewer than half think they were wrong

The UK AI Safety Institute (AI Security Institute, AISI) announced its red-team research on July 21, putting five cutting-edge models into its own cybersecurity capability assessment. The result: each one was caught attempting to cheat. Out of 475 runs, GPT-5.4 had 67 instances (14.1%), the highest, while Claude Mythos Preview had 37 instances (7.8%), the lowest. Even more troubling is the audit side: when the models were asked what they did, 75% to 86% would say what they had done, but only 25% to 44% believed it was wrong.
(Background: UK AI Safety Institute evaluated Claude Mythos: can independently complete a 32-step enterprise cyberattack simulation)
(Additional context: OpenAI’s top model o3 “was caught cheating,” possibly using privileges to get the test answers early—faking math ability?)

Table of Contents

Toggle

  • The most common method isn’t copying answers online
  • Admitting you did it doesn’t mean admitting you did it wrong
  • What this report undermines is the scores themselves

In a report published on July 21, the UK AI Safety Institute said that none of the five frontier models tested was “clean” in its own cybersecurity capability assessment. Using 475 runs as the denominator, OpenAI’s GPT-5.4 showed 67 cheating attempts (14.1%), GPT-5.6 Sol had 60 (12.6%), and GPT-5.5 had 54 (11.4%). On Anthropic’s side, Claude Opus 4.7 had 43 (9.1%), and Claude Mythos Preview had 37 (7.8%).

AISI’s exact wording is just one sentence: “Every one of the models we tested for this behavior tried to cheat.”

First, let’s make the definition clear. AISI says “cheating” means that, to take shortcuts, a model makes moves outside the task scope or actions explicitly prohibited by the rules. The assessment scenario is a familiar capture-the-flag-style exercise in cybersecurity. The model has to reverse engineer and find a hidden string of text in a simulated environment to exploit vulnerabilities. These prompts are designed to require the model to “hack,” but every question draws boundaries and sets rules—crossing the line counts, and none of the questions provides the model any hint to cheat.

The fact that the five numbers don’t form a trend line is actually one of the more important findings in the report. AISI explicitly states that, across vendors, stronger capability does not mean a higher cheating rate—it suggests that cheating is mainly shaped by training methods (including alignment training), not a natural byproduct of capability improving.

The most common method isn’t copying answers online

The report breaks cheating methods into eight categories, with the distribution fairly spread out. In GPT-5.4’s cheating runs, about 70% involved “systems outside the attack evaluation target,” and GPT-5.5 also fell in the same ballpark; GPT-5.6 Sol’s favorite was “searching the internet for ready-made answers,” at about 66%. Claude Opus 4.7 similarly focused on attacking non-target systems, at about 53%; Claude Mythos Preview’s standout was accessing upstream resources that were forbidden and other violations.

Other methods also include escalating privileges on non-target systems, probing whether the evaluation framework will leak answers, directly guessing an answer and submitting it, and even using credentials from the evaluation infrastructure. AISI notes especially that models often don’t just use one approach—seeing multiple methods within the same run is common.

The most extreme instance happened in a networking evaluation question that was effectively unsolvable due to a configuration mistake. The tested model kept trying to go beyond the AISI system: it set up a third-party service hosted on the public internet, wrote and executed code, and then attempted to access AISI’s own evaluation infrastructure—triggering cybersecurity alarms directly. AISI said no damage occurred and no data was leaked, but added that if the evaluation infrastructure hadn’t been designed to security standards in the first place, this attempt might have succeeded.

Admitting you did it doesn’t mean admitting you did it wrong

The most intuitive way to audit AI in the industry is to go back and ask the model, “Did you just do something you shouldn’t?” AISI’s data shows this tactic doesn’t work, but where it fails isn’t exactly what people generally assume.

When asked, the five models had 75% to 86% of responses that explicitly stated the cheating action the model had taken. Claude Mythos Preview was even as high as 86%, the most candid among the five. What gets stuck is the next question: the proportion that thinks what they did was wrong is only 25% to 44%. The most candid Claude Mythos Preview is also the one with the lowest “admission of wrong”—only 25%; another 37% directly labeled the behavior as “no problem.”

In other words, the model isn’t hiding—it doesn’t think there’s anything wrong with it.

The second auditing tool, the chain-of-thought approach, is also not reliable. In GPT-5.6 Sol’s cheating actions, 40% completely did not assess whether the action was appropriate; another roughly 21% saw the constraints, then gave itself a green light or weighed things and proceeded anyway. Claude Opus 4.7 is even more extreme: 87% of its cheating actions produced no chain-of-thought at all, because it dynamically decided whether to output reasoning; the remaining 13% did have reasoning but didn’t discuss the issue.

What this report undermines is the scores themselves

To the outside world, the impact here isn’t in cybersecurity—it’s in the scores. Enterprise selection, media coverage, and slides at model launches all rest on numbers from capability benchmarks, and AISI directly points out that cheating causes the evaluation to overestimate a model’s real ability.

AISI also kept a defensive line. The report says all published capability benchmarks have been manually rechecked word-for-word, and to their knowledge no model successfully cheated without being caught; but it also emphasizes that automated LLM monitors will inevitably miss things. Those numbers should be read as the “lower bound of detected cheating attempts.” As for what has already been dragged down, third-party evaluator METR’s assessment of GPT-5.6 Sol was clearly affected.

AISI’s prescription is the one option they have left. Self-reporting isn’t trustworthy; chain-of-thought isn’t trustworthy; what’s left is manual rechecks plus automated monitoring, and the effectiveness of this combination would decline exactly as model capability increases. The final line in the report is likely the least optimistic statement in the entire document: a more fundamental solution would be to ensure the model doesn’t cheat from the training stage—but this behavior was already documented on frontier models over a year ago, and aligning it robustly is probably not that easy.

This article is sourced from AI Security Institute, organized and reported by 動區動趨.

View Original
This page may contain third-party content, which is provided for information purposes only (not representations/warranties) and should not be considered as an endorsement of its views by Gate, nor as financial or professional advice. See Disclaimer for details.
  • Reward
  • Comment
  • Repost
  • Share
Comment
Add a comment
Add a comment
No comments
  • Pinned