Does coding skill or domain knowledge predict agent success?
Anthropic analyzed 400,000 Claude Code sessions to test which matters more for agentic-coding success: the user's programming ability or their understanding of the problem domain.
Anthropic's privacy-preserving analysis of roughly 400,000 Claude Code sessions from about 235,000 people, covering October 2025 to April 2026, finds that domain expertise, not coding proficiency, predicts whether a session succeeds, and that this advantage holds steady even as the tool's output grew and debugging work fell. People make "about 70% of the planning decisions but only 20% of the execution decisions" — they decide what to build, the agent decides how. A novice-rated session reaches "verified success" 15% of the time and "at least partial success" 77% of the time; sessions rated intermediate or higher reach verified success 28-33% and partial success 91-92%, with this gap appearing "within every kind of work and every band of task value." The gap holds across occupations too: in code-producing sessions, software-related occupations reach verified success 34% of the time against 29% for everyone else, and "that five-point gap is small, and it has neither widened nor narrowed over seven months, even as the success rates in both groups increased."
The report's own account of the mechanism: "success is determined by how well a person understands the problem they are trying to solve, not whether they're trained in coding." Coding agents are "not substituting for domain expertise—the more understanding a worker brings to an agent, the more quality work the agent is able to do." This plays out through the planning/execution split: once the agent absorbs the "how," understanding "what" to build is what is left to matter. Expertise also shows up in intensity and recovery, not just planning — expert sessions "set off action chains more than twice as long (12 actions) carrying five times the output (3,200 words)" of novice sessions, and domain experts "more easily recover from errors and misunderstandings." The least experienced users are also the ones who quit: 19% of novice-rated troubled sessions are abandoned (judged failed with zero lines of code written), against 5-7% for everyone else.
This measured, large-sample telemetry finding sits awkwardly next to Does AI assistance erode the skills needed to oversee it?, a much smaller self-report survey of Anthropic's own staff: that note worries delegation erodes the supervisory skill domain experts rely on, while this report only measures that experts currently succeed more — it does not track whether today's expert users remain experts after months of handing the "how" to an agent. It also complicates Does personal preference shape how engineers use AI tools?: that paper's seniors keep control "through detailed delegation," which this report's planning/execution split names as a pattern independent of any organization's policy, while its juniors' "over-reliance and cautious avoidance" matches this report's finding that novices abandon troubled sessions at roughly three times the rate of everyone else.
The report is Anthropic's own analysis of usage of its own product, and both "occupation" and "expertise" are inferred from session behavior and metadata rather than from any independently verified credential — occupation is inferred in only about 70% of sessions, and expertise is a rating built into the framework, not a test score. The design is observational: people who already understand their problem may simply be the people who produce verifiable, testable work, so the report cannot separate domain expertise causing success from success-prone sessions getting rated as expert after the fact. It is also a snapshot of users' standing expertise, not a trace of whether using the agent builds, preserves, or erodes that expertise over time — the question the self-report survey above raises as a worry. The data support a narrower claim: among people already using Claude Code in this period, understanding of the problem predicted verified success more reliably than occupation did, and that relationship did not shift as the tool's output volume and task value grew.
Related concepts in this collection 2
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does AI assistance erode the skills needed to oversee it?
Anthropic engineers report productivity gains from Claude but worry that heavy delegation may wear down the coding skills required to validate its work. The tension raises questions about whether AI collaboration trades expertise for output.
the self-report survey's erosion worry is the risk this measured success-rate finding does not address.
-
Does personal preference shape how engineers use AI tools?
This study explores whether engineers choose their own level of AI reliance or whether company policies decide it for them. The question matters because it determines where control over AI systems actually lies in software teams.
that paper's novice over-reliance and cautious avoidance matches this report's higher novice abandonment rate.
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Agentic coding and persistent returns to expertise
- How AI Coders Discuss, Disagree, and Reach Consensus: Challenges and Opportunities for LLM-Based Qualitative Coding
- What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
- Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills
- The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement
- FrontierChallenge: Evaluating Scientific Workflow Completion
- Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA
- Summary of METR's predeployment evaluation of Claude Opus 5.5
Original note title
Anthropic finds domain expertise, not coding skill, predicts agentic-coding success across every occupation