Who Should Pace the Frontier? Not Dario Amodei

Paper · Source
Frontier AI Risk & RSI

Source: Dave Karpf, TechPolicy.Press · 2026-09-17

Over the weekend, Anthropic CEO Dario Amodei published a new essay, titled “We Must Pace the Frontier.” Amodei argues for slowing down the pace of AI development, and offers a three-part plan for industry regulation. OpenAI CEO Sam Altman and SpaceX CEO Elon Musk have both responded favorably to his proposal.

Let me start positively. Amodei writes that “A race to the bottom, spurred by commercial incentives, can make [the risks of AI doing serious harm] more acute.” Yes, absolutely.

But it’s also worth noting that Amodei’s company is currently winning that AI race. Anthropic is preparing for its own IPO, boasting about a $30 trillion total addressable market, and scraping together two consecutive quarters of profitability, all while beating OpenAI at its own game. Even if we grant that Amodei’s concern is genuine (and I think it is), it clearly takes a back seat to his commercial ambitions. He is not worried enough to delay going public.

At a high enough level of abstraction, Amodei’s suggestions all sound appealing. He proposes (1) instituting mandatory third-party evaluation and monitoring at all frontier labs, (2) developing a (self-) regulatory regime that all AI labs in democratic countries commit to, and (3) bilateral coordination with China on AI model development. These are all potentially good ideas. But the details matter quite a bit. Amodei’s scheme for pacing the frontier amounts to a series of proposals that all benefit Anthropic. These proposals are all subject to revision if-and-when they get in Anthropic’s way. We should read him accordingly.

And the fundamental problem is this: Dario Amodei believes we ought to trust the AI industry to effectively regulate itself. He has consistently called for regulation, but only of the sort that gives Anthropic a comparative advantage within the industry. He believes the government should mandate the same testing and auditing practices across the industry that Anthropic has already instituted. But, as Lina Khan recently pointed out, the Federal Trade Commission could be much more aggressive in enforcing regulations and holding the industry liable for the harms caused by its products. It is a mistake to let the AI industry shape the contours of its own regulatory system, even if we grant the good intentions of its leaders. We know from history exactly how that will turn out.

  1. Embedded Evaluators. Each frontier AI company commits to giving ongoing, employee-like access to a team of embedded third-party evaluators (such as METR), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes. This is the key step for verifiability of any pacing commitments, and has precedent in the banking industry, which sometimes involves regulatory “supervisors” embedded along with employees. Anthropic is unilaterally committing to this step now. We intend this to be part of a broader push to redouble efforts on our safety and alignment work.

Amodei writes that this “has precedent in the banking industry, which sometimes involves regulatory “supervisors” embedded along with employees.” But the thing about the banking industry is that these regulatory “supervisors” are only there because the government requires them. Without the force of government oversight (and the looming threat of draconian fines), this sort of monitoring doesn’t work.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

What governance mechanisms can effectively constrain widely deployed AI systems? Can AI research automation sustain progress through accelerating feedback loops? Do individually safe AI actions create unsafe outcomes in integrated systems? How can humans maintain effective oversight as AI systems scale? Should governance of agentic AI systems be runtime or design-time? Can models strategically underperform during evaluation to hide capabilities? Does AI assistance help or harm professional skill development?