Toward Measuring AI's Effects on Skill Formation: The Stock-Formation Gap
Large-scale AI deployment data and controlled learning experiments characterize different consequences of the same technology. Deployment telemetry shows that AI use is concentrated in skilled work and frequently supports immediate task performance. It observes tasks, interaction patterns, and outputs, however, not whether users become more capable of performing those tasks independently. Controlled studies measure independent capability more directly, but only in narrower populations and settings, with outcomes that vary substantially by interaction design. We formulate this discrepancy as a stock–formation measurement gap: current systems observe the use of existing expertise more readily than the formation of future expertise. Because formation has historically been society’s recovery mechanism through technological change, the gap matters well beyond any single classroom. We synthesize the experimental and observational evidence by identification strength, use public deployment data as a descriptive illustration of the gap, and identify the missing bridge between interaction traces and unassisted retention and transfer. We then propose a research program that links consented usage records to independent assessments while experimentally varying whether AI supplies answers, hints, feedback, or evaluation. The claim is not that AI has been shown to erode skill formation at population scale. It is that existing measurement cannot determine whether it does, and that this question is both measurable and designable.
Introduction. 1 The Measurement Puzzle Two bodies of evidence describe how people use AI, and they are often read as if they disagree. Deployment telemetry, most extensively the Anthropic Economic Index built from over four million assistant conversations, shows AI use concentrated among skilled professionals, dominated by augmentation-style collaboration (roughly 57% augmentation versus 43% automation), and rewarded by experience (Handa et al. 2025b; Massenkoff et al. 2026). Productivity gains scale with the schooling a task requires (Appel et al. 2026), and field experiments find real performance improvements within the technology’s capability frontier, alongside degraded performance beyond it (Brynjolfsson, Li, and Raymond 2025; Dell’Acqua et al. 2023). Controlled learning experiments, meanwhile, repeatedly find that AI assistance can raise immediate task performance while leaving subsequent unassisted capability flat or reduced, with the outcome varying substantially by how the interaction is designed (Fan et al. 2025; Bastani et al. 2025; Contractor and Reyes 2026). These findings are not in conflict, because they concern different quantities. Telemetry observes tasks, interaction patterns, and immediate outputs. The experiments measure whether a person can later perform without the tool. The first quantity is about the use of existing capability; the second is about the formation of new capability. Reading either literature as an answer to the other’s question produces most of the apparent disagreement in current debates about AI and skills. This paper formulates the discrepancy as a measurement problem, which we call the stock–formation gap: the measurement infrastructure now watching AI adoption observes the deployment of the existing stock of human capital far more readily than the formation of the next generation of it. The central claim is deliberately narrow. The concern is not that current evidence proves AI is eroding learning at scale. The concern is that the dominant measurement infrastructure can register improvements in assisted output while remaining unable to observe changes in independent capability. Controlled studies can measure capability formation but not at deployment scale; deployment telemetry has the scale but does not observe formation; the open question sits where neither instrument reaches. Why the question deserves attention is a matter of history. Across two centuries of technological change, society’s principal recovery mechanism has been education: when technology displaced workers, the next generation adapted by forming the capabilities the new economy demanded. Inequality dynamics track whether the supply of educated workers keeps pace with technological demand (Tinbergen 1975; Goldin and Katz 2008), and skills account for most of the variation in workers’ earnings (Frey 2019). That history rests on an assumption so embedded it is rarely stated: however disruptive a technology, the formation process itself stays intact. Previous technologies automated lower-order cognitive work, and education adapted by moving instruction toward the higherorder work machines could not reach. General-purpose AI is among the first technologies to which learners can, and observably do, hand much of the higher-order work itself; in the largest student dataset, Creating and Analyzing are the most delegated levels (Section 3.3). Whether that delegation erodes formation is an open empirical question; the trials suggest the answer depends on the allocation of cognitive work rather than on the technology as such. The survival of the recovery mechanism should be a measured quantity, not an article of faith. Exposure, meanwhile, is already at scale. Coursework alone accounts for roughly one in eight conversations on a major assistant platform, and directive interaction, in which the user delegates a task and receives an output, is the platform’s most common mode overall (32.6%) (Massenkoff et al. 2026). A parallel economics literature has begun to model the workplace version of the same concern, in which automation of entry-level tasks removes the channel through which juniors acquire tacit skill (Ide 2025; Garicano and Rayo 2025; Afrouzi et al. 2026), with early empirical counterparts in junior hiring declines at AI-adopting firms (Hosseini Maasoum and Lichtinger 2025; Brynjolfsson, Chandar, and Chen 2026). That literature is production-side; the measurement side, to our knowledge, has not been formulated. The ingredients of the stock–formation formulation are individually well established: the performance-learning distinction is classical (Soderstrom and Bjork 2015), cognitive offloading and over-reliance have substantial literatures, and trace-based skill modeling is routine on dedicated tutoring platforms (Section 4.1).
Related work. 3 What the Evidence Establishes We synthesize the evidence in three tiers by identification strength with respect to the assigned condition Z and the realized allocation A. The tiers matter because the literatures use different vocabularies for these variables (unscaffolded versus scaffolded arms in trials are values of Z; directive versus learning-oriented labels in telemetry and externalization versus internalization in the education literature are approximate descriptions of A), and because the strength of causal claims differs sharply across them.
3.1 Tier 1: randomized assignment of the interface (Z) Three experiments randomize the assigned condition itself, and they carry the strongest conclusions. Bastani et al. (2025), in a randomized trial with approximately one thousand highschool mathematics students, found that access to a base GPT-4 tutor raised performance on assisted practice problems by 48% while lowering subsequent unassisted exam scores by 17% relative to no-AI control. A second arm deployed the same model behind pedagogical guardrails that offered hints rather than solutions: assisted practice improved even more (127%), and the scaffolded condition did not show a statistically detectable deficit relative to control, suggesting, but not establishing, that interaction design can mitigate the effect. Because the intervention was a prompting-andguardrail layer rather than a new model, the design lever it isolates is available to any deployer.
Alsaiari et al. (2026), in a semester-long trial (n = 329), randomized the type of AI feedback students received. Directive feedback produced roughly three times higher odds of revision than purely metacognitive prompts (OR = 2.93, p = .009), and hybrid feedback outperformed both (27.5% versus 12.1% revision). The finding cuts against a simple “more learner effort is always better” reading: novices given only reflective prompts often failed to convert them into action, so learner-side load can be overloaded as well as offloaded. Kumar et al. (2026), in a pre-registered trial (n = 968), tested the inverse allocation directly: the participant practiced empathic communication and the AI evaluated and guided, producing a 0.98 SD improvement across most measured dimensions. Together, Tier 1 shows that within these settings the assigned condition Z causally affects ∆K, in both directions. The natural reading, that Z works by shifting the realized allocation A, is plausible but not directly established: these designs randomized interfaces and feedback types, not cognitive behavior, and A was at most partially observed.
Method. 2 The Stock–Formation Framework Seven quantities organize everything that follows. Let S denote the stock state: a person whose capability in a given domain is largely formed and who uses AI to deploy it. Let F denote the formation state: a person whose capability in that domain is still forming. These are domain- and time-relative positions of a person–domain pair, not types of people: a professional learning an unfamiliar library is in formation for that domain (Shen and Tamkin 2026), and a student may be deploying already-formed expertise elsewhere; the Bridge and Mixed entries in Table 2 mark studies that straddle the two, and a continuum refinement is natural. We still speak of stock and formation populations as compositional shorthand: coursework contexts are formation-heavy and professional telemetry is stock-heavy on average, and composition is what determines what deployment data can be taken to show. For any human–AI interaction, let Z denote the assigned condition: the interface, defaults, or configuration under which assistance is delivered, a spectrum running from answer delivery through hints and feedback to evaluation of the person’s own attempt. Let A denote the realized allocation of cognitive work: which parts of the task’s cognitive work the person actually performs in use. The two are not equivalent: a hints interface can be used passively, and an answer-delivery interface can prompt active verification, so Z shapes but does not determine A. Nor is A a single answer-delivery-toevaluation scale; it is best treated as a profile over component functions, for instance planning, generation, monitoring, verification, and correction, and we intend the profile reading wherever measurement is at issue. Let Y denote the immediate output (the essay drafted, the code produced, the problem answered). Finally, let K0 and K1 denote unassisted capability at baseline and after a period of use, measured with the tool removed and including retention over time and transfer to novel problems, with ∆K = K1 −K0. The Z–A distinction parallels a recent reframing in learning analytics of learner agency as the allocation of decision authority across learners, systems, and institutions, where a recurring obstacle is that agency is inferred from behavioral proxies rather than measured directly (Borchers, Viberg, and Kizilcec 2026). The two allocations can vary independently: a learner who retains full decision authority can still delegate the thinking itself, and a system that architects every choice can still leave the cognitive work with the learner. A tracks the cognitive-work allocation, and it inherits the same behavioral-proxy obstacle, which is why the program in Section 5 treats measuring A from traces as a first-class validation problem rather than an assumption. The distinction between Y and ∆K is the learning-science distinction between performance and learning: performance is observable execution during acquisition, while learning is the durable change that survives a delay and a change of conditions, and the two are empirically dissociable in both directions (Soderstrom and Bjork 2015). A second regularity gives the allocation variable its learning-science meaning: conditions that make practice feel harder, the retrieval, monitoring, and error correction that register as struggle, often produce more durable learning than conditions that make performance smooth, the pattern known as desirable difficulties (Bjork, Bjork et al. 2011). On that account, the effortful components of a task are not overhead around learning; they are part of the mechanism of formation, and an assistant that absorbs them by default would absorb formation opportunities along with the friction. That is the mechanistic reason the allocation of cognitive work is the natural moderator, and why effects on ∆K cannot be read off improvements in Y .
Discussion. AI-text detection is unreliable and inequitable in ways that preclude building measurement on it (Corbin et al. 2025), and the automation of grading noted above removes human observation from precisely the point where formation would be noticed (Bent et al. 2025). Privacy-preserving conversation classification demonstrates that interaction traces can be analyzed responsibly at scale (Tamkin et al. 2024); what it currently surfaces is usage, not capability change.
5.5 Governance by design A program linking student conversations to assessment outcomes could itself become an instrument of surveillance or profiling, so its governance must be specified in advance, not retrofitted: participation requires meaningful consent with an opt-out carrying no academic penalty; measurements are controlled by a research function, not by instructors assigning grades or by disciplinary processes; retention is time-limited and labels auditable; classifications are never used for individual high-stakes decisions; and error rates are monitored for disparate impact, particularly for multilingual learners and learners with disabilities, for whom interaction-pattern classifiers are most likely to misread effort. Methods for auditing group-level model bias exist, but typical evaluation samples in educational data mining are underpowered for detecting realistic bias effects (Borchers 2025), so the program’s fairness audits must be powered deliberately rather than run as afterthoughts. Privacy-preserving classification (Tamkin et al. 2024) is a necessary component, not a sufficient one. We regard stakeholder co-design, with students, teachers, and assessment specialists, as a precondition for legitimate deployment of the program, and note that the present paper was developed without it.
5.6 Exploratory: constructs the program may eventually need We flag, explicitly as a hypothesis space rather than a contribution, one further measurement question the program may encounter. The taxonomy the field uses to describe cognitive work (Anderson and Krathwohl 2001) was built when no technology operated at its upper levels. If systems now produce competent output across those levels, the capacities that most distinguish human formation may sit above or beside the named ones: candidates include judgment about which problems are worth solving under genuine uncertainty, and the dispositional constructs the epistemic-cognition literature discusses under epistemic agency and identity, the selfdirected initiation of verification and the stability of one’s justificatory commitments (Greene, Sandoval, and Bråten 2016). Each candidate would require discriminant validation against established constructs before use; sketch operationalizations exist (problem selection under an ill-posed brief; rate of self-initiated verification in think-aloud protocols; stability of commitments under counter-argument), and the think-aloud channel in particular is no longer prohibitive at scale: AI-transcribed think-alouds have been coded automatically for self-regulatory behaviors and related to moment-bymoment performance on tutoring platforms (Borchers et al. 2024b). Whether these constructs add predictive power over the standard taxonomy is itself a testable question the program could carry, and nothing else in the program depends on the answer.
Conclusion. 6.3 Closing The asymmetry of the stakes is why the program should not wait. When technological demand has outrun educational supply in the past, the costs took generations to repair (Goldin and Katz 2008; Frey 2019). If default interaction designs do erode formation, the harm will concentrate among learners with the least scaffolding and mentorship; adoption of learning-supportive tools already skews toward better-supported students in open deployments (Bühler, Bueno, and Kasneci 2026). If they do not, unfounded alarm carries its own costs in access withheld. Measurement is the inexpensive hedge against both errors, and Section 6.1 gave the reason it will not emerge on its own: the feedback channels institutions listen to are fed by the stock’s measured gains, not by the formation side’s unmeasured outcomes. The question this paper formulates is narrow enough to answer and consequential enough to be worth answering: does sustained AI use, under the interaction designs people actually experience, change the formation of independent capability, in which direction, by how much, and for whom? The evidence reviewed here shows the question is real, shows it is currently unanswerable with deployed instruments, and shows that the core technical components of an answer already exist. Until the formation side is measured, decisions about defaults, deployment, and curriculum are being made without the one quantity that determines their long-run consequences. Education has been society’s recovery mechanism through every previous technological disruption; whether it plays that role through this one should not rest on assumption. It is measurable, and this paper has tried to specify exactly how.
Limitations. 6.2 Limitations This synthesis has limits that bound its claims. It is a motivated narrative synthesis, not a systematic review: studies were located by citation chasing and targeted search, the inclusion rule for Table 2 is stated in its note, and no riskof-bias grading was performed; the map is illustrative rather than systematic. The assigned condition Z was randomized in only three studies, and the realized allocation A in none; elsewhere both are observational, and we have marked this accordingly throughout.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Does AI assistance help or harm professional skill development?- Does AI help close skill gaps or preserve them?
- Does AI use during skill-building phases impair how people learn concepts?
- Does AI assistance erode skill development over time among professionals?
- Do quality gains from AI help persist after the tool is removed?
- Can deployment telemetry reveal how expertise forms rather than just how it performs?
- How does automation erode the skills workers need to maintain systems?
- Which professions experience skill erosion versus development with AI tools?
- Can we measure perceived skill change against actual independent task performance?
- Do gains from AI assistance disappear when workers complete tasks alone?
- Can workers retrain faster than AI exposure spreads through occupations?
- How does occupational segregation affect who gains from AI productivity?
- How do institutions shape whether AI enables worker mobility or deepens hierarchy?
- Does AI adoption rise or fall as worker education and wages increase?
- Can persistent agentic workflows predict labor displacement better than task-level exposure?
- Why do firms substitute labor for AI faster than gig worker jobs disappear?
- Can workers move across the divide between technical and non-technical job markets?
- How does AI task concentration within firms affect worker reallocation across jobs?
- Why does AI adoption favor automation over augmentation in female-dominated work?
- Do firms with high AI exposure shed jobs or reshape roles?
- Which occupations show the sharpest gap between AI capability and actual adoption?