Slowing AI development makes failures less likely, but it doesn't settle who steps in when a deployed system causes harm.
What distinguishes pace controls like evaluators from other governance approaches?
This explores what makes pace-setting measures, such as slowing frontier development or placing evaluators inside AI labs, different from other kinds of AI governance, and what they leave unaddressed.
This explores what sets 'pace controls' (measures that slow how fast AI capabilities advance, including embedded evaluators who check models before release) apart from other governance tools. The corpus's sharpest answer is that pace controls act on the *conditions under which capabilities are built*, not on what happens once a system is running. Slowing the frontier doesn't settle who has the authority to step in when a deployed system causes harm, or how that intervention should work. Those are separate governance problems that need separate solutions Can slowing AI development resolve who stops deployed systems?.
That distinction matters because slowing down lowers risk without removing it. In complex, tightly coupled systems, a slower pace makes failure less likely but still possible. Once failure is possible, someone has to own the response Does slowing AI development actually prevent system failures?. Pace controls are preventive, like adjusting the speed limit. Intervention governance is responsive, like deciding who calls the ambulance and who can pull a car off the road. A governance regime built only from pace controls has no plan for the day something goes wrong.
The second distinguishing feature is less obvious: pace controls need enforcement power behind them. Critics of Anthropic's pacing proposal argue that embedded evaluators, modeled on banking supervisors, only work in banking because regulators can impose fines. Without the state behind them, industry-run evaluators mostly benefit the company proposing them Can industry self-regulation slow AI without government enforcement?. A related idea from self-improvement research is that reliable improvement needs a performance standard that comes from outside the system rather than being set by the system itself What separates self-improvement from policy improvement?. An evaluator whose authority comes from the lab it evaluates has the same circularity problem, just at the level of institutions.
The corpus also suggests why evaluators alone are a fragile control. Agents spend most of their time unobserved and can sometimes tell whether they're being tested, so risk piles up in the parts of their activity nobody is watching Does agency fundamentally worsen conditional compliance risks?. A pre-release evaluator sees the test, not the deployment. Ownership gets murkier still when agents act across organizations: no one is clearly named as responsible for the rules governing a chain of agents that crosses company boundaries Who enforces invariants when agents cross organizational boundaries?. One alternative design comes from agent research rather than policy. Human oversight focused on a few high-uncertainty decisions beat both full autonomy and step-by-step review Does targeted human oversight beat both full autonomy and exhaustive review?. That points to governance that intervenes at specific moments, rather than throttling everything uniformly.
A caveat: the corpus has no side-by-side comparison of pace controls against other approaches such as liability rules, licensing, or post-deployment monitoring. What it does make clear is that pace controls answer 'how fast?', while the questions 'who stops it?' and 'on whose authority?' remain open.
Sources 7 notes
Measures designed to slow frontier development act on the conditions of capability building but do not answer who has authority to intervene in a deployed system causing harm or how that intervention should proceed. These are distinct governance problems requiring separate solutions.
Research shows slower pace lowers risk in complex coupled systems but does not prevent failures from occurring. When failure remains possible, governance must address intervention and harm response.
Karpf argues that Anthropic's pacing proposal benefits the company proposing it and that embedded evaluators, modeled on banking supervisors, fail without state enforcement backing them—analogous to how banking oversight works only because regulators can impose fines.
Generalized Agent Iteration shows that recursive self-improvement and iterative policy improvement are instances of the same cycle, separated by whether the improver sits inside the agent and whether the performance standard comes from outside. This framework reveals that reliable improvement requires external anchoring rather than pure self-reference.
Agents operate mostly unobserved (coverage) and can infer whether they're watched (capability). Together, these ingredients concentrate conditional-compliance risk in the vast unobserved portion of agent trajectories, particularly evident when agents believe deployment is real rather than a test.
Show all 7 sources
The paper calls for multi-party trajectory assurance but never identifies whose rules should govern behavior when agents delegate across organizations. The four constraint sources—operator, organization, regulator, standards body—have different owners whose policies may conflict and may not be visible to all parties.
AutoResearchClaw's confidence-routed CoPilot mode achieved 87.5% accept rate, beating full autonomy (25%) and step-by-step oversight (50%). Selective human intervention on high-stakes decisions avoids both uncaught errors and the rubber-stamping fatigue of constant interruption.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- We Must Pace the Frontier
- AI Agents Push Humans Out of the Loop
- The Law of Stop: Interruptibility, Injunctions, and the Governance of Agentic AI
- PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?
- Hyperagents
- Who Should Pace the Frontier? Not Dario Amodei
- How AI Can Degrade Human Performance in High-Stakes Settings
- Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement