Can we measure whether AI erodes independent skill?
Current telemetry tracks how people use AI but not whether they become more capable without it. Existing measurement tools cannot yet determine if AI helps or hurts skill formation at scale.
Toward Measuring AI's Effects on Skill Formation argues that the instruments watching AI use are lopsided. Deployment telemetry observes "tasks, interaction patterns, and outputs" but not "whether users become more capable of performing those tasks independently." The Anthropic Economic Index, built from over four million assistant conversations, shows AI use concentrated among skilled professionals, and the paper's point is what that kind of data cannot see. Controlled learning experiments measure independent capability more directly, but in narrower populations. The authors call the mismatch a stock–formation measurement gap: current systems "observe the use of existing expertise more readily than the formation of future expertise." The claim is deliberately limited: "The claim is not that AI has been shown to erode skill formation at population scale. It is that existing measurement cannot determine whether it does."
The mechanism rests on four of the paper's seven quantities. Z is the assigned condition, the interface or defaults, running from answer delivery through hints to evaluation of one's own attempt. A is the realized allocation of cognitive work, a profile over functions such as planning, generation, monitoring and verification. Y is the immediate output, and ∆K is the change in unassisted capability, measured with the tool removed. Z shapes A without determining it, since a hints interface can be used passively. Because the effortful parts of a task "are part of the mechanism of formation," an assistant that absorbs them by default absorbs formation along with the friction, so effects on ∆K "cannot be read off improvements in Y." The paper's Tier 1 covers three randomized trials and reads Z as moving ∆K in both directions. In one of them, a base GPT-4 tutor raised assisted practice by 48% while lowering unassisted exam scores by 17%, a result Bastani et al. (2025) measured and the paper reports.
The nearest notes are instances of this gap seen from different sides. Does AI assistance help workers learn lasting skills? reports a measured case in which a within-study gain did not carry into later independent work, which is the Y-versus-∆K split in concrete form; this excerpt adds why deployment telemetry would miss that outcome. Can metacognitive feedback stop students from offloading to AI? is an intervention on the handover, the A-side lever, and the paper's program would evaluate it through unaided retention rather than test-day scores alone. The institutional version appears in Do university AI policies actually protect what credentials mean?: boundaries are stated more clearly than the evidence standards that would show what a credential certifies, which is the same measurement failure applied to assessment.
The excerpt establishes less than its framing suggests. It calls itself "a motivated narrative synthesis, not a systematic review," performs no risk-of-bias grading, and notes that the assigned condition was randomized in only three studies and the realized allocation in none. The deployment data is "a descriptive illustration of the gap," and the research program it proposes, which links consented usage records to independent assessments while varying answers, hints, feedback and evaluation, is a design rather than a result. What follows at the strength the evidence allows is narrow: whether sustained AI use changes independent capability is a real, measurable question that deployed instruments cannot currently answer. It does not follow that AI is eroding learning, and the paper itself warns that unfounded alarm carries costs of its own.
Inquiring lines that read this note 10
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Does AI assistance help or harm professional skill development?- Does AI help close skill gaps or preserve them?
- Does AI use during skill-building phases impair how people learn concepts?
- Does AI assistance erode skill development over time among professionals?
- Do quality gains from AI help persist after the tool is removed?
- Can deployment telemetry reveal how expertise forms rather than just how it performs?
- How does automation erode the skills workers need to maintain systems?
- Which professions experience skill erosion versus development with AI tools?
- Can we measure perceived skill change against actual independent task performance?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does AI assistance help workers learn lasting skills?
When workers use generative AI on tasks, do they develop skills they can apply later without AI? This matters because it challenges the assumption that AI-assisted work functions as effective practice.
a measured instance of the performance-versus-capability split this paper says telemetry cannot observe
-
Can metacognitive feedback stop students from offloading to AI?
When learners practice with an AI assistant, does making them aware of the downsides of offloading their work reduce how much they ask the AI to solve for them? And does that change improve their performance on tests without help?
an intervention on the realized allocation A, the lever the paper's measurement program would need to test
-
Do university AI policies actually protect what credentials mean?
Universities are getting better at stating what AI use is allowed, but do their policies explain what evidence proves a student's actual competence? This matters because a credential's value depends on what work the student actually did.
the same gap in institutional evidence standards, applied to credentials rather than learning
-
Can AI narrow the education performance gap?
Does generative AI help lower-education people catch up to higher-education people on complex tasks? This matters because AI's impact on inequality depends on whether it democratizes skills or widens existing gaps.
qualifies: a randomized experiment measured performance after AI removal, finding lower-education users kept part of the gain, so retention is partly measurable beyond telemetry
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Toward Measuring AI's Effects on Skill Formation: The Stock-Formation Gap
- UX Roundup (28 Sep 2026): Bogus Deskilling Research
- How Organizations Use AI: Evidence from ChatGPT
- How AI Impacts Skill Formation
- Gdpval: Evaluating Ai Model Performance On Real-world Economically Valuable Tasks
- GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
- AI Skills Improve Job Prospects: Causal Evidence from a Hiring Experiment
- Who's in Charge? Disempowerment Patterns in Real-World LLM Usage
Original note title
deployment telemetry observes expertise in use but not expertise forming — the stock-formation gap leaves AI's skill effects undetermined