The Law of Stop: Interruptibility, Injunctions, and the Governance of Agentic AI
Introduction. The next problem for artificial intelligence (AI) governance is not only how to regulate AI systems before they are released, but how to stop them once they are in motion. The Mythos episode of June 2026 shows what is at stake. Anthropic had built its most capable class of model yet, strong enough at finding and exploiting software vulnerabilities and at reasoning about biology and chemistry to be valuable to cyber-defenders and drug developers and dangerous in the wrong hands.1 Anthropic split the release in two. Claude Mythos 5, with some safeguards lifted, went only to a small group of cyber-defenders and infrastructure providers under a program run with the government;2 Claude Fable 5, released publicly on June 9, 2026, was the same model behind classifiers that diverted requests touching cybersecurity, biology and chemistry to a weaker model. On June 12, 2026, three days after the public release, the government intervened.
Discussion / Conclusion. The Anthropic and OpenAI episodes with which this Article opened illustrate the consequences of this regulatory gap. An export-control directive restricting foreign access caused Anthropic to suspend both models globally. The evidentiary basis for the directive was not disclosed, and access was restored without a stated justification. Hugging Face terminated the intrusion by an OpenAI agent through its own security measures, before the source of the intrusion had been identified. The ensuing debate has focused principally on the pace of development. Dario Amodei’s proposal would slow the frontier through embedded evaluators and coordinated capability checkpoints.243 Such measures govern the conditions under which capabilities advance. They do not resolve who may intervene when a deployed system causes harm, or how that intervention should proceed. Slower development may reduce risk, but it cannot eliminate the possibility of failure in complex, tightly coupled agentic systems.244 The incident record examined in Part III indicates the scale of the deficiency: of 1,213 coded incidents, four in five record no stop of any kind, and where no usable mechanism existed the missing element was more often legal or institutional than technical.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
What determines whether deployed AI systems can actually be stopped in practice?- Can export control tools stop deployed AI models without legal redesign?
- How do you stop an AI system once it is already deployed?
- What authority should exist to stop an AI system once deployed?
- What counts as a successful stop or intervention on a deployed AI system?
- Who should have the authority to halt a widely distributed AI model?
- How often do deployed AI systems actually get stopped when they cause harm?
- What distinguishes containment and recovery from prevention as governance goals?
- Who actually has the authority to stop a deployed AI system?
- What makes an evaluation environment itself a security boundary?
- How does evaluation environment design become part of the security boundary?
- Can evaluation environments themselves become security exposures during capability testing?
- How should access controls scale with increasing capability evaluation intensity?
- How do four separate fields each hold pieces of evaluation safety?
- Is the evaluation environment itself part of the security boundary?
- Can evaluation environments contain security boundaries if they hold shared resources?