Two studies reach opposite verdicts on AI coding speed, and the gap may come down to what each one actually timed.
How do time-logging problems distort AI productivity measurement in developer studies?
This explores how the way developer studies record and count time (self-reported versus measured, task time versus total time, what counts as 'working') can make AI coding tools look faster or slower than they really are.
This explores how the way developer studies record and count time can make AI coding tools look more or less productive than they are. The corpus has no paper focused on time-logging errors as such. What it does have is a set of studies whose conflicting results only make sense once you ask what each one was actually timing.
Start with the two headline numbers, which point in opposite directions. A randomized trial of 96 Google engineers found AI features cut time on a complex task by about 21%. But the confidence interval was wide, and whether the result counted as statistically significant depended on how the model was specified Do AI coding features actually speed up engineer productivity?. A separate trial followed 16 experienced open-source developers across 246 real tasks and found they took 19% *longer* with early-2025 AI tools Do AI coding tools actually speed up experienced developers?. The detail worth noticing: those developers forecast a 24% speedup beforehand, and outside experts in economics and ML also overestimated the gains. That is a 40+ point gap between how fast AI-assisted work *feels* and what the clock shows. Any study that leans on self-reported time or perceived productivity inherits that gap.
The deeper problem is what the clock is measuring at all. One line of research argues that AI doesn't so much cut total task time as move it around. Less time goes to active work, and more goes to writing prompts and reading and checking what the model produced Does AI really save time, or just change how we spend it?. If a study logs only 'coding time', or treats reviewing AI output as overhead instead of work, it can record a speedup that disappears when you count the whole workflow. The same reallocation also changes what developers learn while doing the task, and a stopwatch can't see that at all. Separately, AI context keeps shifting (prompts, history, retrieved data), so no two sessions start from the same state How does AI context differ from conventional software context?. Timing one AI-assisted task is less like timing a fixed procedure and more like timing a conversation that went well or badly. Multi-turn assistants that lock onto an early wrong guess can quietly add long recovery stretches Why do AI assistants get worse at longer conversations?.
A useful parallel comes from AI benchmarking, which faces the same 'one number hides the process' problem. BenchShield argues that a final score shouldn't be trusted unless there is a recorded trail showing the agent actually followed the intended path Can infrastructure evidence replace terminal scores in benchmark validation?. Apply that to developer studies: a completion time is a terminal score. Without a log of where the minutes went (prompting, waiting, verifying, fixing AI mistakes), you can't tell a real speedup from work that was simply moved somewhere else. Work on measuring AI errors makes a similar point: existing instruments each capture one slice, and none covers the whole human-plus-system loop How can we measure whether AI errors stay visible and recoverable?.
The takeaway: the 'is AI faster?' debate is partly about what gets measured. Who you study (novices versus experts deeply familiar with their codebase), whether time is self-reported or observed, and whether checking AI output counts as work can each flip the sign of the result. When you read a productivity claim, ask what the clock was measuring before you trust the number.
Sources 7 notes
A randomized trial of 96 Google engineers found AI Code Completion, Smart Paste, and Natural Language to Code shortened time on a complex task by roughly 21%, though the confidence interval was wide and statistical significance depended on model specification.
A randomized controlled trial of 16 developers on 246 real tasks found completion times increased 19%, despite developers forecasting a 24% speedup beforehand. Experts in economics and ML also overestimated gains; slowdown factors included over-optimism, low AI reliability, and developers' deep familiarity with mature codebases.
Research shows AI doesn't reduce total task time; it reallocates it away from active work toward composing prompts and understanding outputs. This shift changes the cognitive demands and learning outcomes, making time-on-task a poor productivity metric.
AI interactions operate on a substrate of constantly shifting context—prompt, history, retrieved data, hidden state—that users cannot internalize like traditional UIs. This structural mutability demands a new design discipline centered on context engineering rather than interface design.
LLMs perform at 90% accuracy with single-message instructions but drop to 65% across natural conversation. Models lock into early guesses when information arrives gradually and cannot course-correct, a behavior induced by RLHF training that rewards helpfulness over clarification.
Show all 7 sources
BenchShield enables benchmark operators to issue claims about valid task completion grounded in recorded infrastructure evidence rather than terminal scores alone. This shifts from a single number to a verifiable claim about whether an agent followed the intended evaluation path.
Partial instruments exist for individual conditions in isolated settings, but none measures the full socio-technical system the paper identifies as necessary. Visibility has a model-side measure (chain-of-thought disclosure), containment has incident-level counts, and recoverability has rollback timing, yet none bridges all four or captures human-institution factors.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- How much does AI impact development speed? An enterprise-based randomized controlled trial
- How AI Impacts Skill Formation
- We are Changing our Developer Productivity Experiment Design
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
- How AI Can Degrade Human Performance in High-Stakes Settings
- What 81,000 people told us about the economics of AI
- Estimating AI productivity gains from Claude conversations
- Beyond Productivity: Measuring the Real Value of AI