Does Claude Opus 5.5 fully automate AI research tasks?
METR evaluated whether Opus 5.5 can independently conduct AI R&D work without human oversight. The answer matters for predicting when AI systems might recursively improve themselves without human guidance.
METR's predeployment evaluation of Claude Opus 5.5 — conducted under an unpaid AI R&D assessment agreement with Anthropic, with METR drafting the initial summary, Anthropic given "the opportunity to review and edit the text," and METR signing off on the final text — assesses two separate claims: (A) whether Opus 5.5 would now dramatically accelerate AI R&D at Anthropic, and (B) whether AI already dramatically accelerated the R&D that produced Opus 5.5. On (A), using five capability tasks (Budget NanoGPT Speedrun, LMCA, Gaming Bot, Sunlight, plus trend data from METR's Frontier Risk Report), METR finds Opus 5.5 "improves upon Fable 5.1 across both verifiable tasks ... and harder-to-verify tasks," but is "an incremental improvement ... rather than a discontinuous jump," and is "unlikely to be able to fully automate AI R&D."
METR's reasoning for withholding a full-automation verdict is a skills gap, not a scores gap: full automation would require "large improvements in foresight, prediction, creating one's own feedback loops, and generally other skills that might typically be referred to as researcher 'judgement' or 'taste,'" and "the evidence we have does not suggest that Claude Opus 5.5 represents a large improvement over Fable 5.1 in these 'judgement' skills." On (B), METR is explicit that it did not independently verify the figure it cites for the model's own development: it relays, without endorsing, a "highly experimental and preliminary report from a separate METR assessment of AI R&D acceleration inside Anthropic," which estimated "~1.5X overall acceleration in capabilities due to AI ... with perhaps 30% chance of 2X acceleration." METR's own team states it uses that report's evidence "as an input to our assessment but do not argue directly in defense of its claims."
This sits in tension with Is AI development already being handed to AI systems?, whose cited speedup runs from about 3x to about 52x — far larger than, and independently sourced from, the ~1.5x (30% chance of 2x) that METR relays here secondhand. It also complements Does AI assistance erode the skills needed to oversee it?: METR's judgement-and-taste gap gives a capability-side account of why full delegation stays bounded, alongside that survey's oversight-side account of the same ceiling. And it gives a live instance of the verifiable-versus-harder-to-verify split that Do fixed-budget efficiency gains translate to real research progress? has to bridge from a benchmark score to an R&D trend, here at the point of an actual predeployment evaluation rather than a benchmark proposal.
The excerpt does not establish what rate of continued incremental gains would eventually close the judgement gap, nor whether the ~1.5x figure it relays is itself a reliable measurement — METR flags both as unresolved, noting "the data we have is insufficient for distinguishing consistent, accelerating, or decelerating rates of improvement." The implication, at the strength the evidence allows, is that the two headline risk cases this batch keeps circling — autonomous AI R&D and self-accelerating model development — currently rest more on cited internal estimates and relayed, unverified figures than on capability measurements an outside evaluator could independently confirm.
Inquiring lines that read this note 1
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What human oversight must AI research systems have?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Is AI development already being handed to AI systems?
Anthropic reports rising task length, code authorship, and speedup metrics as evidence that AI systems are taking on development work. The question is whether these measures actually demonstrate autonomous delegation of R&D or reflect improvements in assisted productivity.
contrasts Anthropic's cited 3x-to-52x speedup against METR's unverified, far smaller ~1.5x relayed figure
-
Does AI assistance erode the skills needed to oversee it?
Anthropic engineers report productivity gains from Claude but worry that heavy delegation may wear down the coding skills required to validate its work. The tension raises questions about whether AI collaboration trades expertise for output.
pairs METR's capability-side judgement gap with that survey's oversight-side account of bounded delegation
-
Do fixed-budget efficiency gains translate to real research progress?
The paper measures research efficiency as optimization gains under a fixed evaluation budget, but this differs from the real-world costs of R&D spending and human effort. Does this narrower measurement actually predict whether AI agents reduce the true cost of research discovery?
this evaluation is a live case of the verifiable-to-harder-to-verify bridge that note names
-
How much progress have AI agents actually made on NanoGPT?
METR compares AI agent optimization to human researcher productivity on a popular benchmark. By measuring where their improvement curves intersect, they ask whether autonomous systems are meaningfully accelerating AI R&D or mostly chasing noise.
Evidence for A: expenditure-horizon metric shows NanoGPT agentic optimization yields minimal AI R&D progress, consistent with Opus 5.5's incremental gains
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Summary of METR's predeployment evaluation of Claude Opus 5.5
- Anthropic Risk Report: August 2026 (redacted)
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- UK AISI Alignment Evaluation Case-Study
- Alignment is not solved but it increasingly looks solvable
- RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
- Import AI 475: Swarm scaling; Google DeepMind watermarks biology; and the AI science economy
- Autonomous Research Agents: A Survey of AI Scientists and the Verification Gap
Original note title
METR finds Claude Opus 5.5 is an incremental AI R&D improvement over Fable 5.1 — judgement, not raw capability, is what's missing for full automation