How fast is autonomous AI cyber capability advancing?

Paper · Source
Frontier AI Risk & RSI

Source: UK AI Security Institute · 2026-05-13

The length of tasks frontier models can autonomously complete in our narrow cyber suite has been doubling every few months. This doubling rate has become faster over time, and recent models exceeded our previous trends.

In February 2026, we internally estimated that the length of cyber tasks AI models could complete had doubled every 4.7 months since late 2024 – already an acceleration from our November 2025 estimate of 8 months. Since then, AISI reported on two new models, Claude Mythos Preview and GPT-5.5, which substantially exceeded both doubling rate trends. It is unclear whether this represents a new, faster trend.

Time horizon benchmarks track the length of tasks AI models can complete, measured against the time human experts would take on those same tasks. They are inexact predictors of performance; AI struggles with some tasks humans do quickly, and easily completes others that humans find hard. However, we use this type of benchmark because it offers a measure of AI autonomy from which we can draw trends.

At AISI, we assign each task in our narrow cyber suite an estimate for how long it would take a cyber expert to complete.1 Tasks in our narrow suite require models to identify and exploit cybersecurity weaknesses in target systems, testing skills such as reverse engineering and web exploitation in self-contained setups. These tasks cover only some of the capabilities relevant to real-world cyberattacks.

We deliberately constrain our setup to only 2.5M tokens per task to make results comparable over time. This understates what frontier models can do. We discuss this decision further below.

In February 2026, we estimated that frontier models’ 80%-reliability cyber time horizon had doubled every 4.7 months since reasoning models emerged in late 2024, given a 2.5M token limit. This was around half our November 2025 doubling time estimate, which was 8 months for both 50% and 80% reliability. Claude Mythos Preview and GPT-5.5 have since significantly outperformed this trend. At the time of writing, it's unclear whether Mythos Preview and GPT-5.5 represent an isolated break from existing rates of progress or are part of a new, faster trend.

Mythos Preview and GPT-5.5 have large upper-bound error bars due to near-100% success rates on our narrow cyber suite’s longest tasks, even with the 2.5M token limit.2 Our tasks are also not long enough to determine how sharply the models’ reliability would deteriorate at higher task lengths. This places some of the latest models at the limit of what our narrow test suite can measure.

We have also observed further evidence of cyber autonomy beyond our narrow task suite. AISI’s cyber ranges (shown below) measure AI models’ ability to complete cyberattacks against small, undefended enterprise networks, where initial access has already been gained. Each cyber range requires sustained planning and execution capability; more detail on them can be found in our recent paper.

In AISI’s latest testing, the newer Mythos Preview checkpoint completed both our cyber ranges, solving the range “The Last Ones” in 6 of 10 attempts and the previously unsolved “Cooling Tower” in 3 of 10 attempts. This was the first time that a model completed the second of our two cyber ranges. GPT-5.5 solved “The Last Ones” on 3 of 10 attempts.

Lines of inquiry this paper opens 10

Research framings built by reading the notes related to this paper — the questions it feeds into.

Can AI research automation sustain progress through accelerating feedback loops? How should we measure frontier AI models' cyber exploitation capabilities? How can defenders detect and contain coordinated agent attacks? What limits recursive self-improvement in autonomous AI systems? How do evaluation environment design choices affect AI security? Can AI systems evade safety evaluations through reasoning manipulation? What human oversight must AI research systems have?