Does generative AI narrow education-based productivity gaps? Evidence from a randomized experiment

Paper · arXiv 2608.04198 · Published August 4, 2026
Social Theory and Society

Does generative artificial intelligence (AI) reinforce or reduce productivity differences across workers? Existing evidence largely studies AI within firms and occupations, where organizational selection compresses educational heterogeneity, leaving unclear whether AI narrows productivity gaps across individuals with different levels of education. A related question is whether the productivity gains from AI reflect productive use of the tool or mere delegation, fading once AI is unavailable, and whether the answer differs across education groups. We address both questions using a randomized online experiment outside firms, in which 1,174 adults aged 25–45 complete an incentivized, workplacestyle business problem-solving task with or without a generative-AI assistant, followed by a non-AI-assisted follow-up module. AI increases performance for all participants, with substantially larger gains for lower-education individuals. In the control group without AI, higher-education participants outperform lower-education participants by 0.548 standard deviations; with AI, this gap falls to 0.139 standard deviations, closing about three-quarters of the initial gap.

Introduction. Technological change has long been a driver of inequality in labor markets. Historically, innovations have often disproportionately benefited highly educated workers, contributing to widening gaps in employment and earnings (Acemoglu and Autor, 2011; Goldin and Katz, 2008). Generative artificial intelligence (AI) now presents a new paradigm: more directly than many previous technologies, these systems can perform complex cognitive tasks. This raises the possibility that generative AI could be “skill-democratizing” rather than skill-biased, as suggested by Autor (2024), enabling workers with limited formal training to perform tasks previously requiring extensive education and thus empowering workers across the education distribution. Acemoglu (2024), by contrast, cautions that AI’s effects on labor markets may be more modest and unequally distributed than these early discussions suggest. Whether AI weakens the link between formal education and productivity or instead reproduces existing skill advantages in a new form remains an open empirical question.

Discussion / Conclusion. This paper studies whether generative artificial intelligence reinforces or reduces productivity differences across individuals with substantially different levels of formal education. Using a randomized online experiment conducted outside firms, we obtain four main findings. First, on whether AI narrows the education-based performance gap, we find that AI closes about three-quarters of the baseline gap, with larger gains for lower-education participants. Second, on the nature of these gains, the follow-up results do not support a pure temporary-delegation interpretation: treated participants do not perform worse than controls once AI is removed, and lower-education participants retain part of their gain, although a sizable education gap remains. Third, on when gains carry over, intensive AI assistance predicts strong main-task performance, but follow-up performance is substantially higher when assistance is combined with sustained task engagement.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How does AI adoption across firms reshape employment and inequality? Does AI assistance promote real skill development or substitute for independent learning? What happens to knowledge when intelligence becomes tokenized like a commodity? How should designers communicate what AI systems truly are and can do? Can inoculation prompting prevent emergent misalignment after reward hacking? How well do AI systems understand human social norms? How does persona conditioning amplify demographic stereotyping and bias in models? What determines whether deployed AI systems can actually be stopped in practice?