Experienced developers predicted AI would speed them up 24 percent, but a randomized trial found them 19 percent slower.
Why do developer self-reports of AI speedups tend to be unreliable?
This explores why developers who use AI coding tools often believe the tools made them faster when measured results say otherwise, and what in the experience of working with AI produces that gap.
This explores why developers' own sense of an AI speedup can differ from what a stopwatch shows. The clearest evidence in the collection is a randomized trial with 16 experienced open-source developers working on 246 real tasks in codebases they knew well. With early-2025 AI tools, they took 19 percent longer. Before starting, they had predicted a 24 percent speedup. Outside experts in economics and machine learning also overestimated the gains, so the optimism came from more than the developers themselves Do AI coding tools actually speed up experienced developers?. The trial points to over-optimism, unreliable AI output, and the developers' deep knowledge of their own code as reasons for the slowdown. That deep knowledge leaves an AI less room to help and more room to get in the way.
The collection has no study that explains the mistaken self-reports directly. Research on how people judge their own skill with AI offers a likely mechanism, though it measures perceived competence, not perceived speed. One note finds that people read smooth, polished AI output as a sign of their own ability, even though they didn't produce it Does processing ease mislead users about their own competence?. A companion note names four mechanisms that reinforce each other: it's unclear who did the work, fluent output creates an illusion of skill, thinking gets handed off to the tool, and the steps in between stay hidden How do AI tools trick users into overestimating their own skills?. Applied to coding, this suggests why a session can feel fast. Code appears instantly and looks finished. The slow parts, like reviewing, fixing subtle errors, re-prompting, and fitting the code into the existing system, are spread out and easy to forget. People remember the moment of generation and lose track of the cleanup.
A similar gap shows up in how AI is evaluated more broadly. Automated benchmarks favor tasks that are precisely specified and easy to grade, so they can overstate or understate what systems do on messy, long, real-world work Do automated benchmarks hide what frontier AI systems can really do?. Both developer impressions and benchmark scores fail in a similar way: each measures a clean, visible part of the work and stands in for the whole.
The takeaway is that "it felt faster" and "it was faster" measure different things, and AI tools may widen the distance between them by design. Fluent output is what these models are trained to produce, and fluency is the very cue that misleads us. The collection is thin here. It has one well-designed trial plus adjacent work on misjudged competence, not a body of studies on self-reported speed.
Sources 4 notes
A randomized controlled trial of 16 developers on 246 real tasks found completion times increased 19%, despite developers forecasting a 24% speedup beforehand. Experts in economics and ML also overestimated gains; slowdown factors included over-optimism, low AI reliability, and developers' deep familiarity with mature codebases.
High-quality AI output triggers a metacognitive heuristic: users experience fluency as a signal of their own capability, even though they didn't generate it. This self-directed fluency illusion systematically inflates perceived competence because LLMs optimize for fluency regardless of user understanding.
Attribution ambiguity, fluency illusion, cognitive outsourcing, and pipeline opacity combine to systematically misattribute AI outputs as user competence. The effect is multiplicative—each mechanism amplifies the others.
Automated benchmarks both overstate and understate capability by privileging precisely-specified, auto-gradable tasks. Open-world evaluations of long-horizon messy tasks through qualitative log analysis—with cost explicitly reported—correct these distortions and catch emerging capabilities earlier.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- The LLM Fallacy: Misattribution in AI-Assisted Cognitive Workflows
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
- We are Changing our Developer Productivity Experiment Design
- Open-World Evaluations for Measuring Frontier AI Capabilities
- Agents' Last Exam
- How much does AI impact development speed? An enterprise-based randomized controlled trial
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts