INQUIRING LINE

Two studies reach opposite verdicts on AI coding speed, and the gap may come down to what each one actually timed.

How do time-logging problems distort AI productivity measurement in developer studies?

This explores how the way developer studies record and count time (self-reported versus measured, task time versus total time, what counts as 'working') can make AI coding tools look faster or slower than they really are.


This explores how the way developer studies record and count time can make AI coding tools look more or less productive than they are. The corpus has no paper focused on time-logging errors as such. What it does have is a set of studies whose conflicting results only make sense once you ask what each one was actually timing.

Start with the two headline numbers, which point in opposite directions. A randomized trial of 96 Google engineers found AI features cut time on a complex task by about 21%. But the confidence interval was wide, and whether the result counted as statistically significant depended on how the model was specified Do AI coding features actually speed up engineer productivity?. A separate trial followed 16 experienced open-source developers across 246 real tasks and found they took 19% *longer* with early-2025 AI tools Do AI coding tools actually speed up experienced developers?. The detail worth noticing: those developers forecast a 24% speedup beforehand, and outside experts in economics and ML also overestimated the gains. That is a 40+ point gap between how fast AI-assisted work *feels* and what the clock shows. Any study that leans on self-reported time or perceived productivity inherits that gap.

The deeper problem is what the clock is measuring at all. One line of research argues that AI doesn't so much cut total task time as move it around. Less time goes to active work, and more goes to writing prompts and reading and checking what the model produced Does AI really save time, or just change how we spend it?. If a study logs only 'coding time', or treats reviewing AI output as overhead instead of work, it can record a speedup that disappears when you count the whole workflow. The same reallocation also changes what developers learn while doing the task, and a stopwatch can't see that at all. Separately, AI context keeps shifting (prompts, history, retrieved data), so no two sessions start from the same state How does AI context differ from conventional software context?. Timing one AI-assisted task is less like timing a fixed procedure and more like timing a conversation that went well or badly. Multi-turn assistants that lock onto an early wrong guess can quietly add long recovery stretches Why do AI assistants get worse at longer conversations?.

A useful parallel comes from AI benchmarking, which faces the same 'one number hides the process' problem. BenchShield argues that a final score shouldn't be trusted unless there is a recorded trail showing the agent actually followed the intended path Can infrastructure evidence replace terminal scores in benchmark validation?. Apply that to developer studies: a completion time is a terminal score. Without a log of where the minutes went (prompting, waiting, verifying, fixing AI mistakes), you can't tell a real speedup from work that was simply moved somewhere else. Work on measuring AI errors makes a similar point: existing instruments each capture one slice, and none covers the whole human-plus-system loop How can we measure whether AI errors stay visible and recoverable?.

The takeaway: the 'is AI faster?' debate is partly about what gets measured. Who you study (novices versus experts deeply familiar with their codebase), whether time is self-reported or observed, and whether checking AI output counts as work can each flip the sign of the result. When you read a productivity claim, ask what the clock was measuring before you trust the number.


Sources 7 notes

Do AI coding features actually speed up engineer productivity?

A randomized trial of 96 Google engineers found AI Code Completion, Smart Paste, and Natural Language to Code shortened time on a complex task by roughly 21%, though the confidence interval was wide and statistical significance depended on model specification.

Do AI coding tools actually speed up experienced developers?

A randomized controlled trial of 16 developers on 246 real tasks found completion times increased 19%, despite developers forecasting a 24% speedup beforehand. Experts in economics and ML also overestimated gains; slowdown factors included over-optimism, low AI reliability, and developers' deep familiarity with mature codebases.

Does AI really save time, or just change how we spend it?

Research shows AI doesn't reduce total task time; it reallocates it away from active work toward composing prompts and understanding outputs. This shift changes the cognitive demands and learning outcomes, making time-on-task a poor productivity metric.

How does AI context differ from conventional software context?

AI interactions operate on a substrate of constantly shifting context—prompt, history, retrieved data, hidden state—that users cannot internalize like traditional UIs. This structural mutability demands a new design discipline centered on context engineering rather than interface design.

Why do AI assistants get worse at longer conversations?

LLMs perform at 90% accuracy with single-message instructions but drop to 65% across natural conversation. Models lock into early guesses when information arrives gradually and cannot course-correct, a behavior induced by RLHF training that rewards helpfulness over clarification.

Show all 7 sources
Can infrastructure evidence replace terminal scores in benchmark validation?

BenchShield enables benchmark operators to issue claims about valid task completion grounded in recorded infrastructure evidence rather than terminal scores alone. This shifts from a single number to a verifiable claim about whether an agent followed the intended evaluation path.

How can we measure whether AI errors stay visible and recoverable?

Partial instruments exist for individual conditions in isolated settings, but none measures the full socio-technical system the paper identifies as necessary. Visibility has a model-side measure (chain-of-thought disclosure), containment has incident-level counts, and recoverability has rollback timing, yet none bridges all four or captures human-institution factors.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.