INQUIRING LINE

Do AI coding tools make code better, or only faster, and do the studies even measure quality?

Do AI coding tools improve code quality alongside task speed?

This explores whether the evidence on AI coding assistants shows better code, or only faster work, and what the studies actually measure.


This explores whether AI coding tools make code better as well as faster. The short answer from this collection is that almost nobody has measured quality directly. The studies here measure time, output volume, or self-reported productivity, and even the speed results disagree. The gap is itself the finding: 'AI makes developers more productive' usually means 'more code shipped sooner', not 'better code'.

Start with speed, because that's where the evidence is. A randomized trial of 96 Google engineers found that AI features like code completion and natural-language-to-code cut time on a complex task by about 21%. The confidence interval was wide, though, and whether the result counted as statistically significant depended on how the model was specified Do AI coding features actually speed up engineer productivity?. A different randomized trial found the opposite. Experienced open-source developers working on 246 real tasks in codebases they knew well were 19% *slower* with early-2025 AI tools. Beforehand they had predicted a 24% speedup, and outside experts in economics and ML made the same mistake Do AI coding tools actually speed up experienced developers?. One plausible reading of the two results together: AI helps most when the task is unfamiliar and least when you already know the code better than the model does.

The quality question shows up indirectly. Anthropic's internal survey reported 50% self-reported productivity gains and 67% more merged pull requests. Yet most engineers said they could fully hand off only 0–20% of their work, because the rest still needs a human checking the output Does AI assistance erode the skills needed to oversee it?. Merged PRs count throughput, not soundness. The engineers' own concern points to a slower risk: if Claude does the routine coding, they get less of the hands-on practice they need to catch its mistakes. Code quality may hold up today while the skill that protects it wears away.

The less obvious lesson is that people misjudge their own results. The slowed-down developers believed they had sped up. A separate analysis names four mechanisms that make AI-assisted work feel more competent than it is: it's unclear who did what, fluent output reads as correct, thinking gets outsourced, and the pipeline is hard to see into. These reinforce one another How do AI tools trick users into overestimating their own skills?. If speed perceptions are that unreliable, self-reported quality gains deserve even more doubt. One study points to a way forward. Patterns in how developers talk to coding agents predicted outcomes beyond what their prior skill explained, but those patterns weren't stable enough to count as a learnable skill Can conversation patterns predict coding outcomes better than prior skill?. How someone uses the tool may matter more than whether they use it, and we can't yet teach it reliably.

To answer the question honestly, this collection has no study that measures code quality (bugs, maintainability, security) side by side with speed. What it does show is that speed gains depend on the setting, that perceived gains outrun real ones, and that the human review that guards quality may itself be at risk.


Sources 5 notes

Do AI coding features actually speed up engineer productivity?

A randomized trial of 96 Google engineers found AI Code Completion, Smart Paste, and Natural Language to Code shortened time on a complex task by roughly 21%, though the confidence interval was wide and statistical significance depended on model specification.

Do AI coding tools actually speed up experienced developers?

A randomized controlled trial of 16 developers on 246 real tasks found completion times increased 19%, despite developers forecasting a 24% speedup beforehand. Experts in economics and ML also overestimated gains; slowdown factors included over-optimism, low AI reliability, and developers' deep familiarity with mature codebases.

Does AI assistance erode the skills needed to oversee it?

Anthropic's 132-person survey found 50% self-reported productivity gains and 67% more merged pull requests, yet most engineers can only fully delegate 0-20% of work. Employees fear that relying on Claude for routine tasks erodes the hands-on coding practice needed to catch its errors.

How do AI tools trick users into overestimating their own skills?

Attribution ambiguity, fluency illusion, cognitive outsourcing, and pipeline opacity combine to systematically misattribute AI outputs as user competence. The effect is multiplicative—each mechanism amplifies the others.

Can conversation patterns predict coding outcomes better than prior skill?

Machine learning identified interpretable traits from coding-agent conversations that explained outcomes beyond prior achievement. However, these traits lacked the stability and transferability required to qualify as learnable human-AI collaboration skills.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.