How fast is AI accelerating its own development inside labs?
Anthropic proposes measuring how much AI systems now do the work of building themselves, and publishes initial metrics claiming AI leads 26% of its R&D work. The question matters because AI-driven development could speed up capability gains while making human oversight harder.
Anthropic argues that the public needs visibility into how fast frontier AI development is moving inside the labs building it, and it publishes a first snapshot of what that visibility could look like. The starting premise is that AI systems "have begun to automate more of the process of building themselves." The snapshot gives three figures. "Claude 'leads' 26% of Anthropic's AI R&D work," and the share of work "at or above 'AI collaborates'" is "above 90%." As of August 2026, "approximately 30,000 agents" were doing research and engineering work at any one time on Anthropic's most-used internal platform. The excerpt says those agent figures cover that platform only.
The reasoning is that labs increasingly use AI to build future models, which lets them "develop more capable models more quickly" and test them more before release, but that models accelerating their own development "could make it more challenging for humans to understand or control these systems." Measurement is meant to show "how close the world is to reaching recursive self improvement (a model fully autonomously building its successor)." The excerpt sets out three measurement areas: the share of R&D work along a path "from 'AI collaborates' toward 'AI leads'"; the oversight of semi-autonomous agents that delegate to one another; and compute allocation, which Anthropic calls "among the most verifiable inputs to the AI R&D process" and a possible "critical lever in a future pacing effort." Anthropic also says the numbers "would shift if there were coordination on pacing the frontier, as called for by Anthropic CEO Dario Amodei," and it plans to embed independent third-party evaluators with access comparable to internal risk assessment teams.
The nearest notes approach the same question from other angles. The Atria Dawn analysis reads task records to see who proposes and who decides, and Anthropic's scale asks a similar who-leads question at lab level, but as a self-reported share rather than task-level data. ASI-Bench's staged withdrawal of guidance tests how far an agent can go unaided on a benchmark, while Anthropic's scale tracks a comparable progression inside a working lab. The ASI-Evolve loop of analysis and reuse is the kind of AI-driven research cycle a lab-wide share would have to count, though the excerpt does not say whether such loops appear in the counted work. The efficiency bridge in the AI R&D benchmark note uses a different unit altogether: research efficiency under an evaluation budget, not a share of work.
The excerpt does not establish how "leads" or "collaborates" are assigned, whether shares are counted by time, tasks or tokens, or what the methodological Appendix says, because that Appendix is not in the excerpt. The 26% and above-90% figures are Anthropic's own internal numbers, and the independent evaluators it plans to embed are described as a future step. The excerpt gives one date for the agent count and no time series, so it cannot show whether the share is rising. The monitors that constrain the agents are "summarized below," but that summary is not in the excerpt, so how they constrain the agents is unstated. What follows is a self-reported snapshot from one developer on one platform. By the source's own account, it would need third-party verification and regular publication before it could show how the pace is changing.
Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
How much guidance do AI systems need to conduct research independently?
ASI-Bench tests whether AI can explore open-ended research problems by progressively removing human methodological guidance. This matters because existing benchmarks cannot distinguish between AI that follows instructions well and AI that can autonomously discover and verify new knowledge.
same autonomy progression, staged on a benchmark rather than observed in a working lab
-
How should AI agents and humans divide research tasks?
In building its own foundation model, Atria Dawn studied how to split work between agents and human researchers. Understanding this division matters for designing effective human-AI collaboration in technical R&D.
the same who-leads question, answered from task records rather than a self-reported lab-wide share
-
Can AI research itself without losing human oversight?
Explores whether AI systems can internalize the human judgment and insight-distillation that normally drives research progress, and what this means for maintaining meaningful human control over AI advancement.
a closed AI research loop of the kind a lab-wide share would have to count
-
Do fixed-budget efficiency gains translate to real research progress?
The paper measures research efficiency as optimization gains under a fixed evaluation budget, but this differs from the real-world costs of R&D spending and human effort. Does this narrower measurement actually predict whether AI agents reduce the true cost of research discovery?
a different unit: research efficiency under an evaluation budget, not share of work
-
Is AI development already being handed to AI systems?
Anthropic reports rising task length, code authorship, and speedup metrics as evidence that AI systems are taking on development work. The question is whether these measures actually demonstrate autonomous delegation of R&D or reflect improvements in assisted productivity.
Evidence for: Anthropic cites Claude's share of merged code and a 3x-to-52x speedup as evidence AI development is being handed to AI
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Measurements for understanding the pace of AI development inside frontier labs
- How AI is transforming work at Anthropic
- The Economics of Recursive Self-Improvement
- Summary of METR's predeployment evaluation of Claude Opus 5.5
- Introducing Anthropic Interviewer: What 1,250 professionals told us about working with AI
- Anthropic Economic Index report: Uneven geographic and enterprise AI adoption
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- When AI builds itself
Original note title
Anthropic argues the public needs visibility into frontier-lab AI R&D pace — its snapshot reports Claude leads 26% of Anthropic's AI R&D work