Does removing AI tools actually measure real skill loss?
Nielsen questions whether lab experiments that take away AI and measure performance decline tell us anything useful about how AI affects workers in real jobs where the tool never gets removed.
Nielsen's complaint targets a specific experiment and the genre it belongs to. Shang Wu and colleagues (University of California, Irvine; HCOMP 2026) had 124 participants solve logic puzzles across three phases — no AI, optional AI, no AI again — and found that "everybody improved with practice, but AI users improved less, and their assisted scores overestimated their later unassisted scores." Nielsen calls this design "the confiscation study: hand people a tool ... let them use it, take it away, measure the decline, and conclude that the tool erodes skill." He grants "the measurement is valid," but argues "it rests on a false premise: outside the research lab, nobody confiscates the tool."
His reasoning runs by analogy: "We don't evaluate carpenters by confiscating their hammers and saws at the end of their apprenticeships and timing how long they take to drive a nail with a rock... The carpenter keeps the hammer. The knowledge worker keeps the AI." Because removal never happens in practice, he argues the field has been measuring the wrong dependent variable, and calls for a different program: "study what happens to users' higher-level skills after AI takes over the lower-level ones. Assume everybody has AI. Never measure performance without it... And distinguish among the ways people use AI, because their effects on learning are far from equal." He credits the Wu study with one finding that survives the design's flaw — skill gains tracked time spent reasoning before asking for help, not request frequency — but insists "no Bayesian model, however elegant, rescues a study that measures the wrong outcome."
This is a methodological argument, not a new measurement, and it sharpens Can we measure whether AI erodes independent skill? from the opposite direction: that paper's "stock-formation gap" says deployment telemetry can't see skill formation because it only watches assisted use with the tool left in place; Nielsen's complaint is that the lab-side alternative, removing the tool to test retention, watches a condition nobody actually lives in. Between the two, almost no clean window remains onto skill formation under permanent AI use. Can metacognitive feedback stop students from offloading to AI? is closer to the study Nielsen's program calls for: the assistant stays available throughout, and the intervention targets how answers get requested rather than whether access exists; its finding that heavy answer-requesters scored worse lines up with Nielsen's own point that help-seeking frequency, not help-seeking access, is what should be tracked. Do junior developers choose AI based on their ability to verify results? supplies the kind of distinction among usage modes that Nielsen says confiscation studies collapse into a single "used AI" variable.
The excerpt is commentary, not a study: Nielsen names the design flaw and proposes a direction but runs no experiment of his own, cites no body of confiscation-study results that reverse once the tool stays available, and does not specify which higher-level skills to measure or how. The implication he draws — that most published deskilling findings describe a condition absent from deployed use — is only as strong as the premise that the confiscation design can't be patched, a premise he asserts rather than demonstrates. What the excerpt does establish plainly is a standard worth applying to any deskilling claim encountered afterward: ask whether the tool was removed to produce the result.
Inquiring lines that read this note 5
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Do AI coding tools measurably improve developer productivity and code quality? Does AI assistance help or harm professional skill development? How can AI systems reliably guide voters without introducing political bias? Does AI-assisted work increase total productivity or just shift time?Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can we measure whether AI erodes independent skill?
Current telemetry tracks how people use AI but not whether they become more capable without it. Existing measurement tools cannot yet determine if AI helps or hurts skill formation at scale.
makes the opposite complaint: telemetry can't see skill formation because it never removes the tool at all
-
Can metacognitive feedback stop students from offloading to AI?
When learners practice with an AI assistant, does making them aware of the downsides of offloading their work reduce how much they ask the AI to solve for them? And does that change improve their performance on tests without help?
the kind of study Nielsen calls for: tool stays available, intervention targets how it's used
-
Do junior developers choose AI based on their ability to verify results?
Can junior developers reliably decide when to use AI by assessing whether they can check the output themselves? This matters because it reveals how newcomers self-regulate AI use amid pressure to adopt it quickly.
evidence for distinguishing ways people use AI, the distinction Nielsen says confiscation studies collapse
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- UX Roundup (28 Sep 2026): Bogus Deskilling Research
- Toward Measuring AI's Effects on Skill Formation: The Stock-Formation Gap
- Stranded Credentials: Keeping Online Reputation Systems Informative in the AI Era
- How AI Impacts Skill Formation
- AI Skills Improve Job Prospects: Causal Evidence from a Hiring Experiment
- Gdpval: Evaluating Ai Model Performance On Real-world Economically Valuable Tasks
- GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
- Microsoft New Future of Work Report 2025
Original note title
Nielsen argues the confiscation study measures the wrong outcome because nobody confiscates AI outside the lab