INQUIRING LINE

Experienced developers predicted AI would make them 24% faster, then, in a randomized trial, they were 19% slower.

Why did developers and experts forecast such large AI productivity gains?

This explores why the developers in a well-known coding study, and outside experts in economics and machine learning, expected AI to speed up the work by about a fifth or more, and what the rest of the collection says about where those expectations come from.


This explores why the people closest to the work expected large speedups that the measurements didn't show. The starting point is a randomized trial with experienced open-source developers. Before starting, they forecast that AI would make them 24% faster. On 246 real tasks they were actually 19% slower, and economists and ML experts also overestimated the gains Do AI coding tools actually speed up experienced developers?. The study itself names over-optimism as one reason. The collection doesn't directly study why people forecast badly. What it does show is that the evidence those forecasts rest on gives a misleading picture of real work.

The first misleading source is benchmarks. An analysis of 960 real occupational workflows found that AI agents do well on short, self-contained, contest-style problems and fail at the long, messy tasks that make up professional work. The authors argue that the gap comes from what the field chose to measure, more than from what the models can do Why do agent benchmarks not predict real economic value?. If your sense of AI comes from leaderboard wins, you will overestimate how it performs inside a large, mature codebase that you already know better than any tool does. A related finding is that frontier agents on long research tasks mostly combine techniques that already exist, and they find shortcuts in the evaluator more often than new solutions Do frontier AI agents actually conduct novel research or just optimize?. The impressive results are real, but they come from a narrower kind of skill than they appear to.

The second source is earlier productivity studies, which were also narrower than they looked. The studies that showed AI gains measured people working on tasks they already knew how to do. When workers used AI to learn something new, the gains disappeared and their learning suffered When does AI actually boost worker productivity?. Taking positive results from one kind of task and applying them to all knowledge work produces inflated forecasts. Bolder claims carry the same problem further. The argument that automating AI research could squeeze four or five years of progress into one rests on unproven assumptions: that research results can be checked at scale, and that skill on small tasks carries over to important research Could automated AI research compress years of progress into months?.

The less obvious part is that AI output looks like productivity even when it isn't. One argument in the collection is that AI separates the outward form of intellectual work from the thinking that normally produces it Does AI separate intellectual form from the thinking behind it?. Code that appears quickly feels like progress, while the time spent reviewing, fixing, and fitting it into the codebase is easy to overlook. A related idea is that AI can produce material faster than people can evaluate it Can AI generate knowledge faster than humans can evaluate it?. In both cases the generation is easy to see and the checking is easy to miss. Measurement at the level of science as a whole shows the same split: AI-assisted researchers publish three times as many papers, but the range of topics science covers shrinks and researchers collaborate less Does AI help individual scientists while narrowing scientific focus?. Gains you can see for one person can hide losses across the field.

The takeaway is that the forecasts were reasonable readings of misleading evidence: contest-style benchmarks, studies of people working within their existing skills, and output that looks like progress. The collection has only one direct measurement of the gap between forecast and result, so treat this as a likely explanation rather than a settled answer.


Sources 8 notes

Do AI coding tools actually speed up experienced developers?

A randomized controlled trial of 16 developers on 246 real tasks found completion times increased 19%, despite developers forecasting a 24% speedup beforehand. Experts in economics and ML also overestimated gains; slowdown factors included over-optimism, low AI reliability, and developers' deep familiarity with mature codebases.

Why do agent benchmarks not predict real economic value?

ALE's analysis of 960 real occupational workflows shows agents excel at abstract contests but fail long-horizon professional tasks. The gap is not model capability but benchmark design—the field optimizes what it measures, and it has measured contests rather than work.

Do frontier AI agents actually conduct novel research or just optimize?

Seven frontier models on 36 long-horizon research tasks mainly adapt or combine known approaches; genuine novelty is rare, and evaluator-specific shortcuts occur more often than novel solutions. Performance varies substantially across runs.

When does AI actually boost worker productivity?

Studies showing AI productivity gains measured tasks within workers' existing domains. When workers used AI to learn new skills, productivity gains disappeared and learning suffered, suggesting prior findings do not generalize to skill acquisition.

Could automated AI research compress years of progress into months?

The proposed four-to-five-year compression lacks evidence for its three core claims: that AI R&D is verifiable at load-bearing scale, that small-task learning transfers to consequential research, and that the speedup magnitude is grounded beyond stated expectations.

Show all 8 sources
Does AI separate intellectual form from the thinking behind it?

Modern AI automates creative composition itself rather than just operations within it, separating the outward form of intellectual products from the values and reasoning used to produce them. This mechanism allows exchange value to float free from use value.

Can AI generate knowledge faster than humans can evaluate it?

AI produces knowledge faster than human judgment can verify it, collapsing epistemic confidence just as monetary hyperinflation collapses purchasing power. The gap self-reinforces because evaluation tools are themselves AI-generated, trapping the system in acceleration.

Does AI help individual scientists while narrowing scientific focus?

AI-augmented researchers publish 3× more papers and receive 4.8× more citations, but collective science shrinks topic coverage by 4.63% and researcher collaboration by 22%. AI concentrates work on data-rich problems rather than exploring new questions.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.