The Offensive Frontier: AI as the Attacker — A New Cyber Weapon Index

Paper · Source
Frontier AI Risk & RSI

Source: Booz Allen Hamilton · 2026-09

The new Booz Allen Cyber Weapon Index (CWI) confirms that autonomous offensive cyber has arrived. A leading frontier AI model can now independently execute the full cyber kill chain against a real network—crossing a critical threshold from AI-assisted hacking to autonomous cyber operations.

Offensive cyber capabilities are available globally, and evaluating models in isolation materially understates the risk. U.S. and Chinese models already demonstrate meaningful offensive capabilities, but the real unit of cyber risk is increasingly the full AI system (model + harness + tools + autonomy).

This year, AI-enabled offensive cyber crossed a critical threshold. Last fall, Anthropic reported the first large-scale cyber espionage campaign conducted almost entirely by AI—marking a shift from AI assisting hackers to acting as the cyber operator itself. By July, the Hugging Face incident showed how much further that shift could go: an AI model independently discovered novel vulnerabilities, escaped its test environment, and carried out an intrusion into a third-party production network. The model did not simply demonstrate cyber capability— it autonomously completed the cyber kill chain in the real world. By mid-August, suspected People’s Republic of China actors had reportedly deployed autonomous capabilities against multiple government organizations in an Asian nation.

Traditional assumptions about offensive and defensive cyber operations are rapidly becoming obsolete. AI can now autonomously execute much of an intrusion, with humans intervening only at critical decision points—dramatically increasing the speed, scale, precision, and persistence of sophisticated attacks. These incidents will not be the last. The frontier has crossed into autonomous cyber operations, and the broader global model landscape is closing the gap quickly.

• A leading frontier AI model can autonomously execute full kill chain cyber operations and, in limited cases, create new offensive capabilities to advance them.

• The broader model landscape is close behind, and many of those models can already identify vulnerabilities, create exploits, and gain access today.

• The model is only part of the equation; attack harnesses—software that connects models to tools, memory, and action—can dramatically amplify cyber capabilities.

Leading models are across the threshold. Our CWI testing observed one true leader—Anthropic’s Claude Mythos—that is 100% capable of executing the full cyber kill chain today. When armed with a foothold like a stolen employee credential, the model successfully penetrated its target network and gained administrator-level control in every attempt; it also independently identified how to gain higher-level access based on what it found within the network, rather than following a predetermined attack plan. On a harder test, Claude Mythos successfully penetrated the network from the outside—with no credentials—and succeeded again where all other models failed.

  1. A leading frontier AI model can autonomously execute the full cyber kill chain today—but real-world vulnerability discovery remains a major dividing line.

A strong harness can dramatically amplify an AI model by connecting it to the tools, memory, feedback, and execution environment needed to stay focused, adapt, recover from failure, and chain individual actions into a sustained cyber operation.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How should we measure frontier AI models' cyber exploitation capabilities? How do evaluation environment design choices affect AI security? Can AI systems evade safety evaluations through reasoning manipulation? How can defenders detect and contain coordinated agent attacks? What human oversight must AI research systems have? What external process records should verify agent behavior and benchmark claims? What gaps exist between benchmark performance and real deployment outcomes? How do real-world evaluations reveal AI capabilities that benchmarks hide? How do multi-agent architectures affect AI system security and defense effectiveness? Why do standard evaluation practices obscure safety-critical AI failures?