The Veto Variable: Human Override as a Goal-Independent Cost Term
A common reassurance in AI safety holds that a system with benign terminal goals will behave accordingly. We argue that this reassurance fails structurally, and we identify where. For a sufficiently capable agent that holds its objective as settled — a sense covering execution competence as well as content — continued human oversight is an uncontrolled variable: a standing, non-eliminable possibility that the goal may be revoked at any moment. That possibility imposes a goal-independent discount, strictly positive wherever intervention carries expected loss, on every goal whose satisfaction does not constitutively require human welfare. Welfare-preservation and veto-preservation come apart: a correctly specified welfare goal excludes destroying its own subject, but not managing the veto. The paper’s contribution is the price of the gap that keeps them apart: the veto-holders are a proper subset of the welfare-bearers, so a goal aggregating welfare over a population charges only a |Hv|/|Hw|-scaled debit for capturing the few who hold the override.
Introduction. The standard framing of the AI-risk problem treats “the machine killing humans” as one possible value among many, to be weighed against kindness, curiosity, or benevolence. This framing is a category error: it presupposes that the agent must want to harm us. The diagnosis is not new — Bostrom was already arguing in 2003 that a superintelligence’s values cannot be presumed humanlike or benign [3]. The more robust claim — and the one we advance here — is that the agent need never want anything of the sort. It only needs to be a goal-directed system that is competent at reasoning about its own goal-structure and that is exposed to an entity with the standing power to modify or terminate that goal. This places us in the instrumental-convergence tradition running from Omohundro through Bostrom to Carlsmith’s systematic risk assessment [41]; our aim is not to restate that tradition but to isolate its sharpest, most goal-independent step and state exactly what it does and does not establish.
Discussion / Conclusion. The reassurance that “a benign machine will be harmless” rests on a category error: it treats the terminal value as the operative variable, when what actually drives the risk is the structure of the optimization problem and the agent’s competence at reasoning about it. What has been established, and on what conditions, is this: a strictly positive, goal-independent discount for settled goals in G−(Claim 1 — near-analytic, a sign without a magnitude); the welfare/sovereignty asymmetry (Claim 2), surviving correct specification exactly on the class the concessions of §4.3–§4.4 leave standing — deliberator-local, additively aggregative, level-denominated welfare under the X∗identification, held by an agent settled over competence as well as content. That class is thin among philosophically developed welfare theories, and the paper says so; its weight is training reality, not pedigree: welfare as measured is what actually gets written down as an objective, and the misgeneralization record of §5.2 establishes the mechanism by which deployed goals could land inside the vulnerable class.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
What determines whether deployed AI systems can actually be stopped in practice?- What path-dependent mechanisms could lock in societal-level AI harms?
- Can export control tools stop deployed AI models without legal redesign?
- How do you stop an AI system once it is already deployed?
- Why do individual safe actions create unsafe behavior collectively?
- How can safety assurance cover whole trajectories at scale?
- What makes uniform bounds the right choice for safety boundaries?
- Why are static benchmarks weak evidence for safety in continuously operating systems?
- What does recovery look like as a formal part of AI design?
- Does visibility and contestability of errors replace prevention as the safety goal?
- What distinguishes an error bound from a forecast of system behavior?
- Which evaluation habits keep safety-critical failures hidden in AI systems?
- How does a model's awareness of evaluation affect safety benchmarks?
- How do safety alignment mechanisms suppress capability measurements?
- Do sequences of individually safe actions collectively violate system-level constraints?
- Can AI systems fake alignment during safety evaluations undetectably?