SYNTHESIS NOTE
Topics›Frontier AI Risk & RSI›this note

How fast is AI cyber autonomy actually advancing?

The UK AI Security Institute measures how long autonomous tasks frontier models can complete, finding doubling every few months. But whether recent models signal a fundamentally faster trend remains unclear.

Synthesis note · 2026-10-06 · sourced from Frontier AI Risk & RSI

The UK AI Security Institute reports that the length of tasks frontier models can autonomously complete in its narrow cyber suite "has been doubling every few months," and that this doubling rate "has become faster over time." Its February 2026 internal estimate put the doubling time at 4.7 months since late 2024, already faster than the 8 months in its November 2025 estimate. Claude Mythos Preview and GPT-5.5 then "substantially exceeded both doubling rate trends," and AISI says it is "unclear whether this represents a new, faster trend." AISI also cites cyber autonomy beyond the narrow suite. On its cyber ranges, which test attacks against small, undefended enterprise networks where initial access has already been gained, the Mythos Preview checkpoint solved "The Last Ones" in 6 of 10 attempts and the previously unsolved "Cooling Tower" in 3 of 10, the first time a model completed that second range. GPT-5.5 solved "The Last Ones" in 3 of 10.

The measure is a time-horizon benchmark. Each task carries an estimate of how long a cyber expert would take, and model performance is compared against that human time. AISI calls such benchmarks "inexact predictors of performance," since AI "struggles with some tasks humans do quickly" and "easily completes others that humans find hard," but it uses them "because it offers a measure of AI autonomy from which we can draw trends." The narrow suite asks models to identify and exploit weaknesses in self-contained targets, testing skills such as reverse engineering and web exploitation, and the excerpt says these "cover only some of the capabilities relevant to real-world cyberattacks." The 2.5M-token cap per task is set "to make results comparable over time," a choice AISI says "understates what frontier models can do."

Three neighbors bear on this. The Do cybersecurity benchmarks actually measure exploitation? note argues that exploitation, where a vulnerability becomes an attack, is the under-evaluated stage of cybersecurity benchmarking. AISI's suite is built around that same identify-and-exploit step, so the two sources read as complementary rather than in conflict; the excerpt does not compare coverage, so it cannot say which measure reaches further. The time horizon is also a single axis, which is the case the Does a single benchmark score actually predict agent readiness? note makes against single-axis benchmarks. AISI's caveats echo that note, but AISI uses the axis for trend-tracking, not readiness claims, so the two do not conflict. The How soon do AI researchers expect artificial general intelligence? note aggregates what researchers forecast about timelines; this source reports a measured trend in one domain and does not connect it to those forecasts.

The excerpt does not establish that the acceleration is a new trend. AISI says it cannot yet tell whether Mythos Preview and GPT-5.5 are "an isolated break from existing rates of progress" or part of a faster one. The top-end estimates are weak: both models reach near-100% success on the suite's longest tasks, which gives "large upper-bound error bars," and the tasks are too short to show how reliability falls at higher lengths, which "places some of the latest models at the limit of what our narrow test suite can measure." The excerpt also gives no method or interval for the 4.7-month estimate. The supported claim is narrower than a forecast: AISI's own estimates have sped up, the newest models sit above that trend, and the slope for those models is poorly pinned down. Because the token cap understates capability by AISI's account, the reported horizons are conservative readings, not projections of where the trend goes.

Inquiring lines that read this note 5

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can AI research automation sustain progress through accelerating feedback loops? How should we measure frontier AI models' cyber exploitation capabilities? How can defenders detect and contain coordinated agent attacks? What limits recursive self-improvement in autonomous AI systems?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 114 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

UK AI Security Institute finds the cyber time horizon doubling every few months and accelerating — whether recent models mark a new trend is unclear