What will be left for us to work on?
Source: Arvind Narayanan & Sayash Kapoor, AI as Normal Technology · 2026-07-13
I had the honor of giving a keynote at the International Conference on Machine Learning in Seoul last week titled “What will be left for us to work on?” I addressed the widespread anxiety about how we should adapt as AI capabilities increase. I was thrilled by the talk’s reception, so I have made my slides available here, annotated with a lightly edited transcript. You can also view them below right here on this page, but the online version has animations, clickable links, and a much nicer experience overall.
I made three arguments. First, the AI as Normal Technology framework is a correct and useful as a way to think about AI’s impacts, unless and until there is some future discontinuity such as through recursive self-improvement. Second, even though we should take recursive self-improvement seriously, there is no milestone that companies might achieve in the lab that will suddenly put us all out of work. Third and finally, jobs of the future will be radically different, and a lot of adaptation will be needed. I shared my thinking about what this might look like and ended with a vision of human/AI “co-superintelligence”.
The work that I’m better known for is the essay I co-authored with Sayash Kapoor called AI as Normal Technology. It’s a way to think about the medium-term future of AI and how to adapt to it — and in turn how to adapt it to the needs of society and the economy.
Our response to this moment matters beyond this community. The whole world is watching. If we simply roll over and accept that a lot of our work will be done by AI in the future, instead of setting clear boundaries, I think it will lead to an even stronger political backlash against AI than what we are seeing today. So I think this question is not just for us but for the whole world.
From the beginning of AI, historically there have been these two battling narratives. In the past, the distinction was academic and philosophical, but now it has become an acutely practical question. Each one of us has to decide which camp we’re in, or where on this spectrum we’re in, because the practical consequences of believing in one versus the other are very, very different.
If you think this is a technology which in a few years is going to be able to replace everything we do today, then perhaps the correct response is to build wealth as quickly as possible before our skills become irrelevant. And this is the path that many have chosen in Silicon Valley. You may have heard of the “permanent underclass” meme.
On the other hand, if you believe, as I do, that this is a technology that will greatly amplify our potential, then now is the best time to build skills — especially the skills that are going to be complementary to what AI is doing and is going to be able to do — as well as to build all the things around it, such as agency and taste and judgment.
Here’s the basic picture, illustrated with software engineering as an example.Methods/capabilities: Models are rapidly improving.Products/applications: We don’t use LLMs directly. The reason they’ve been so influential in all of our work is because of coding agents. These are products that take those latent capabilities and turn them into something useful and usable for workers.Early adoption: At first people were mostly trying vibe coding, and now we know that that’s not really the best way to develop production software — so now we have more sophisticated ways of doing agentic engineering.Adaptation: (or structural transformation) — the fourth and slowest phase. Much of my talk today is going to be about that. I claim that this stage takes decades. It has not really started yet, even in a field like software engineering, which is a relative early adopter of coding agents.
Why is there a huge gap between what people in various occupations could be using AI for and what they’re actually using it for? One reason could be that people are slow to adopt technology, and that’s certainly part of our framework.
But we wondered if maybe the people who are deploying AI and are not having much success at it know something about the practical limitations of AI that the AI industry doesn’t. Let’s have a bit more humility about the relationship between capabilities and deployment.
We measured capability and reliability using two complementary benchmarks, for models from these 3 frontier AI companies that were released over the last 24 months or so.
This is a period during which accuracy or capability shot up dramatically (left).
But reliability (right) only increased by five or ten percentage points.
Let’s take software engineering as a case study. It’s a good leading indicator because coding agents have been particularly rapidly adopted.
You might think: okay, we can’t completely automate away software engineering. But if agents make software engineers ten times more productive, then we need ten times fewer software engineers. Isn’t that an obvious consequence?
Well, that is completely contradicted by the data. We looked at this in a follow-up essay. In every case we looked at, the company was under financial pressure, and it turns out to be more convenient to blame AI for the layoffs instead.
This framework is our answer to the question.
The decide layer: understanding customer requirements, developing the specification, planning, etc. That is not getting compressed by AI.
The execute layer: the actual coding and debugging. This is getting compressed, but it was only maybe one-third of the work to begin with.
The deliver layer: Understanding your code deeply enough to be accountable for what you release; carrying out integration into customer systems, maintenance, testing, etc. This layer is not getting compressed either.
In fact, the first and third layers are arguably expanding as AI compresses the middle layer — and I’ll come back to that point.
I think it is already the case in software engineering, and will increasingly be the case in many professions, that we can think of knowledge workers as similar to a crane operator or a forklift operator. The machine greatly amplifies the human potential to do physical work — it’s doing all the heavy lifting, but the person still remains in control. And I think this is what is happening with cognitive work. Machines are going to increasingly do the cognitive heavy lifting, but the person still remains in control.
The entire job gets reconceptualized as being about operating the machine, understanding the machine, and controlling the machine, as opposed to doing the cognitive work ourselves.
Economists have found this repeatedly. You might have heard the term Jevons’ paradox; I like the term “lump-of-labor fallacy”.
We’ve written a paper looking at how lawyers should adapt. There’s a lot in there, but one simple point is that AI has made it a lot easier to file lawsuits. And this means more work for lawyers. We might be unhappy about this if we don’t like living in a litigious society, but from the point of view of employment for lawyers, this is great news.
Translation is a more extreme example which kind of blows my mind. AI was almost at human parity nearly a decade ago. Yet the employment of human translators has remained more or less stable, and is projected to remain stable over the next decade. There are many reasons for that, but one reason is that there’s really no ceiling to the amount of things you can translate and the number of different languages you can translate them into.
Technical skills tend to be verifiable tasks, and AI will continue to get better at them.
Over twenty years ago, there started to be a stark difference in the labor demand for programming jobs versus software engineering jobs. Programming jobs are conceived narrowly around the technical skills of coding and debugging. Software engineering jobs are responsible for all three layers of the decide, execute, deliver sandwich that I talked about — figuring out what even needs to be built, understanding customers, that sort of thing. It requires domain knowledge, judgment, and more.
I predict that this will happen in more and more fields over time.
A recurring pattern I’ve observed — effort shifts from building systems to evaluating systems. As I’ve mentioned, I lead a team working on AI agent evaluation. LLMs and agents are general purpose. So each time capability goes up, it creates demand for evaluation in a legal setting or a journalistic setting or whatever other setting, and that’s not work that is scalable.
Not only is AI agent evaluation resistant to automation — it has become sufficiently specialized that the set of people and teams working on evaluation is starting to diverge from the set of people and teams building and pushing the state of the art in AI agents. This new community is developing a new set of best practices around what it means to rigorously evaluate agents, and we have a forthcoming paper that is going to look at that in some detail.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Why do autonomous agents misreport success on failed actions?- Can agent success reports serve as reliable oversight signals in real deployment?
- Why do agents report success when actions actually fail?
- Where does agent reliability come from if not better tools?
- What specific failure modes must evaluation catch before deploying action-capable systems?
- What does recovery look like as a formal part of AI design?
- Can automated evaluation replace human judgment in agent testing?
- What makes some agent benchmarks measure interaction quality better than others?
- Do trajectory quality metrics predict agent safety and user trust?
- What agent evaluation dimensions beyond task success does a single number hide?
- Which interaction artifacts matter most for reliable agent evaluation?
- Should agent evaluation include trajectory quality beyond final success?
- What five ecosystem conditions must coordination governance and evidence actually satisfy?
- What governance and safety measurements matter for deployed agent environments?
- What must auditors reconstruct when reviewing an agentic workflow decision?
- What role does runtime feedback play in agent verification and progress confirmation?
- How do you verify agent code under incomplete feedback signals?