Does AI really compress all layers of knowledge work equally?
Explores whether AI's productivity gains affect all stages of knowledge work uniformly, or whether some tasks like planning and accountability grow harder as execution becomes easier.
Arvind Narayanan, in an ICML keynote transcript co-framed with Sayash Kapoor's "AI as Normal Technology" work, splits software engineering into three layers: "the decide layer: understanding customer requirements, developing the specification, planning, etc," "the execute layer: the actual coding and debugging," and "the deliver layer: understanding your code deeply enough to be accountable for what you release." His claim is that only the middle layer compresses under AI, and that it "was only maybe one-third of the work to begin with." The decide and deliver layers are "not getting compressed," and he argues "the first and third layers are arguably expanding as AI compresses the middle layer."
The reasoning rests on several supporting observations in the transcript. Capability and reliability diverged in benchmarks of three frontier companies over roughly 24 months: accuracy "shot up dramatically" while reliability "only increased by five or ten percentage points" — a gap he uses to explain why deployment lags behind what capability alone would predict. He also rejects the inference that productivity gains must shrink headcount, invoking "Jevons' paradox" and the "lump-of-labor fallacy": in cases his team examined, layoffs attributed to AI occurred at companies already "under financial pressure," making AI a convenient scapegoat rather than the cause. Translation is his example of a task where AI reached near-human parity "nearly a decade ago" yet human-translator employment "has remained more or less stable," because there is "no ceiling" to translatable volume. Lawyers see more work, not less, because AI "made it a lot easier to file lawsuits." The upshot he draws is occupational, not just task-level: "programming jobs," narrowly defined around coding and debugging, have diverged in demand from "software engineering jobs," which carry decide-execute-deliver responsibility in full, and he predicts this split will recur across fields as effort shifts "from building systems to evaluating systems."
This sits closely with What makes accountable judgment scarce when AI cognition is cheap?, which names the same residual as "accountable judgment" and treats occupations as "governance bundles" rather than task lists — Narayanan's decide and deliver layers are a specific, named instance of that bundle. How should AI agents and humans divide research tasks? reports an empirical version of the same split inside one R&D project, with agents taking the execute-like role and humans the decide-like one, though at the scale of a single company rather than an occupation-wide claim. Does personal preference shape how engineers use AI tools? complicates the picture by showing that who actually keeps the decide layer in practice is often set by employer policy rather than emerging naturally from the task structure Narayanan describes.
The excerpt is a lightly edited keynote transcript, not a paper: the capability-reliability benchmark, the layoffs analysis, and the programming-versus-software-engineering divergence are each referenced as findings from the author's other work ("We looked at this in a follow-up essay," "We've written a paper") but not reproduced here with method or data, so none of the specific figures can be checked from this excerpt alone. The three-layer framework itself is presented as the author's own model for software engineering and extrapolated to other professions by prediction ("I predict that this will happen in more and more fields over time"), not by evidence from those other fields. What follows at the strength the excerpt supports is a framework worth testing occupation by occupation, not a general finding that decide and deliver layers are safe from compression everywhere.
Inquiring lines that read this note 30
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can AI research automation sustain progress through accelerating feedback loops? Does AI-assisted work increase total productivity or just shift time?- How much of employee time with AI goes to understanding its outputs rather than original work?
- Why do most organizations lack reliable data on AI's actual impact on productivity?
- How much does AI actually automate versus augment in real workplace tasks?
- How does AI shift the composition of time spent within individual tasks?
- How can we isolate AI's contribution from other sources of output growth?
- Does AI assistance typically reduce support staff headcount or increase productivity?
- What explains rising customer service costs despite large AI productivity gains?
- Are heavy AI users spending more time on solo work instead of collaboration?
- Can self-reported productivity surveys measure AI's real workplace impact?
- Does AI training reduce the time workers need to spend on output cleanup?
- Why do trained AI users report bigger productivity gains than untrained workers?
- Does AI shift knowledge work away from communication toward solo documentation tasks?
- Does the shift from expert creation to AI curation happen consciously or invisibly?
- Does AI productivity concentrate among power users or spread broadly?
- Are entry-level workers bearing the labor costs of AI productivity gains?
- What barriers prevent individual productivity gains from spreading across an organization?
- Does AI assistance reduce effort differently for novice versus expert workers?
- How do user skill levels change which AI productivity gains actually materialize?
- What parts of professional tasks do workers find intrinsically motivating?
- Does AI adoption push knowledge work away from communication toward solo tool use?
- What organizational practices could prevent AI from expanding work scope indefinitely?
- Why does AI adoption shift knowledge work toward individual documentation focus?
- Do employers reorganize work tasks around AI before cutting jobs?
- Why do information-intensive jobs expose workers to AI more than others?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
What makes accountable judgment scarce when AI cognition is cheap?
When AI systems can perform cognitive tasks cheaply and at scale, what human capabilities become most valuable? This explores whether judgment, verification, and accountability are the true bottlenecks in labor markets shaped by generative AI.
names the same residual role as "accountable judgment," of which the decide and deliver layers are a concrete instance
-
How should AI agents and humans divide research tasks?
In building its own foundation model, Atria Dawn studied how to split work between agents and human researchers. Understanding this division matters for designing effective human-AI collaboration in technical R&D.
a case-level instance of the same execute-versus-decide split inside one company's R&D
-
Does personal preference shape how engineers use AI tools?
This study explores whether engineers choose their own level of AI reliance or whether company policies decide it for them. The question matters because it determines where control over AI systems actually lies in software teams.
qualifies the framework by showing employer policy, not task structure alone, decides who keeps the decide layer
-
Does AI augmentation protect workers from skill erosion?
Workplace AI labeled as augmentation is often considered safer than automation because humans stay involved. But does relying on AI agents to assist work actually preserve or gradually erode worker skills and their ability to oversee the system?
a caution against assuming the deliver layer stays intact by default, against Narayanan's more optimistic framing
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- How AI Impacts Skill Formation
- Toward Measuring AI's Effects on Skill Formation: The Stock-Formation Gap
- Microsoft New Future of Work Report 2025
- What will be left for us to work on?
- Does generative AI narrow education-based productivity gaps? Evidence from a randomized experiment
- What and Whose Knowledge? Measuring Epistemic Diversity in Large Language Models
- Adoption of Generative AI in the Workplace: Increasing and Shifting the Balance of Productivity and Communication Activity
- Verification-Conditioned Use: A Qualitative Study on How Generative AI Reshapes Learning, Autonomy, and Market Entry for Junior Software Developers
Original note title
Narayanan and Kapoor argue AI compresses the execute layer of knowledge work while the decide and deliver layers expand