Anthropic’s latest model release, Mythos Preview, has been positioned as a step change in AI capability for cybersecurity-related tasks. Unlike earlier frontier models, it has been described as particularly strong at handling structured, multi-stage technical problems that resemble real-world intrusion scenarios.
That framing alone was enough to get tech leaders like Microsoft, Apple, and AWS on board the Glasswing project - where Mythos is being tested - with the closed use of the model generating rave reviews from the companies.
Yet such fervor surrounding the project has also triggered interest from security researchers and policymakers. The UK government’s AI Security Institute (AISI) has announced it has run an independent evaluation of the model, with the aim of understanding whether Mythos genuinely represents a new class of cyber capability, or whether it simply sits in line with recent frontier systems from companies like Anthropic and other leading AI developers.
Early findings suggest a more nuanced picture than the hype might imply.
What Was Actually Tested
The evaluation from the UK’s AI Security Institute builds on a long-running benchmarking program that uses Capture the Flag (CTF) challenges to measure AI systems against cybersecurity-style problems. These tests range from relatively simple exploitation exercises through to more complex, multi-step scenarios designed to mimic real intrusion workflows.
According to AISI, Mythos Preview reaches a new high point on its “Apprentice” level tasks, completing more than 85% of challenges in that category. However, this performance is broadly comparable to other frontier models, including GPT-5.4, Opus 4.6, and Codex 5.3, which all sit within a similar range of accuracy across multiple difficulty tiers.
While Mythos does not dramatically outperform its peers on isolated tasks, its behavior across extended sequences of actions raises more interesting questions about how far AI systems are progressing in practical offensive cyber capability.
The more significant part of the evaluation focuses on a test environment called “The Last Ones” (TLO). This simulation was designed to model a 32-step data extraction attack across a fictional corporate network. It requires chaining together multiple actions across different systems, mirroring the kind of sustained effort that would typically take a skilled human operator many hours to complete.
In this environment, Mythos showed clearer differentiation. It was the first model to fully complete the TLO challenge end to end, although only in a minority of attempts. On average, it completed 22 out of 32 steps per run, compared with a lower baseline from earlier models such as Claude 4.6, which averaged around 16 steps. AISI also noted, however, that Mythos still struggles with more advanced scenarios such as the “Cooling Tower” test, which simulates disruption of industrial control systems.
Breakthrough or Incremental Progress?
On paper, Mythos does not represent a dramatic leap in raw cyber task performance when compared to other leading models. On isolated tasks, it is broadly aligned with systems like GPT-5.4 and Anthropic’s own recent releases.




