SYNTHESIS NOTE
Topics›Domain Specialization›this note

Did developers opt out of METR's AI study because of selection bias?

METR's August 2025 developer productivity study may have missed its most AI-dependent workers. Understanding whether selection effects—developers refusing to work without AI tools—distorted the speedup estimates matters for interpreting what the data actually shows about AI's real-world impact.

Synthesis note · 2026-10-06 · sourced from Domain Specialization

METR's February 2026 update says its second developer productivity experiment no longer gives a reliable reading of how much AI tools speed up experienced open-source work. The study began in August 2025 with 57 developers, and each task was assigned to an "AI allowed" or "AI disallowed" condition. The raw estimates were a speedup of -18% for the 10 original developers who rejoined (confidence interval -38% to +9%) and -4% for the 47 newly recruited developers (confidence interval -15% to +9%). The earlier study had found AI use made tasks "19% longer," with a confidence interval of +2% to +39%. METR believes developers are "likely" more sped up now than in early 2025, but calls its own data "only very weak evidence" for the size of that increase, and both new intervals include zero.

The mechanism METR gives is selection rather than sampling noise. Agentic tools such as Claude Code and Codex made some developers unwilling to take part: "an increased share of developers say they would not want to do 50% of their work without AI," even though the study paid $50 an hour, down from $150 in the original study, which METR thinks "also likely contributed." Task selection narrowed as well. Between 30% and 50% of surveyed developers said they had held back tasks they did not want to do without AI, so the study misses "tasks which have high expected uplift." METR lists further problems it judges smaller: developers running several agents at once found per-task time hard to record, output quality differed between conditions, and one developer completed none of the AI-disallowed tasks assigned. METR reads the selection as pushing its central estimate down, so the figure is "a lower-bound," yet the same passage calls it "likely a bad proxy" for the real productivity effect.

The two nearest notes mark the contrast. The Does Figma Make speed up design task completion? trial reports roughly 20 percent shorter completion times under randomization, and METR's excerpt shows what randomization cannot do alone: it balances only the people and tasks that stay in the study, so a clean design can still miss the users most likely to gain. The time-logging problem is the measurement side of the reallocation described in Does AI really save time, or just change how we spend it?, where time moves toward prompting and evaluating output. METR's developers found that shift makes time records unreliable when they worked an unrelated task while an agent ran. METR's planned use of observational data, such as aggregate commit statistics and transcripts, points toward the trace-based approach in Does generative AI shift knowledge workers away from communication?, which avoids self-reported time and task choice but measures activity rather than output.

The excerpt does not establish how large AI's effect on these developers is, or whether it has grown since 2025. METR's growth claim rests on conversations and surveys, and it calls the data "very weak evidence" for that claim. It also says developer self-reports of "very high speedups" "can be quite unreliable," which leaves the study without a usable estimate of either size or direction. The sample is experienced open-source contributors with a median of 10 years' experience, working on their own repositories across 143 repos and 800+ tasks, so it does not describe developers in other settings. The strength the evidence supports is narrow: this randomized design has stopped measuring what it was built to measure, and METR is redesigning it. The implication is that any claim about the size of the speedup needs a design that keeps high-adoption developers in the sample. METR's proposed options, including more intensive experiments, fixed-task experiments and developer-level randomization, aim at that, but the excerpt reports none of them as results.

Inquiring lines that read this note 1

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What human oversight must AI research systems have?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 73 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

METR says selection effects make its August 2025 developer study an unreliable signal of AI speedup — developers opted out rather than work without AI