SYNTHESIS NOTE
Topics›Reasoning o1 o3 Search›this note

Can distance alone rank which substrates resist reward hacking?

Does the amount a method changes a model reliably predict how exposed it is to evaluator errors? The paper tests whether a single distance metric can universally order vulnerability across weights, selection, and prompts.

Synthesis note · 2026-09-24 · sourced from Reasoning o1 o3 Search

The abstract states the formal core: "We formalize a distance-dependent upper bound on evaluator disagreement and a capacity ordering for nested policy classes, then show why distance alone cannot establish a universal ranking of vulnerability. An exact finite-output illustration demonstrates how the location of a scoring defect changes the behavior favored by each method." The discussion separates the pieces: "A distance-dependent error envelope gives an upper bound; class inclusion gives a capacity ordering; the geometry of accessible behaviors and the effectiveness of search determine what a particular system actually finds."

Three kinds of statement. The envelope bounds how much evaluator disagreement can matter given how far a policy moves; "how far a policy moves" is the excerpt's only gloss on distance. The ordering says that when one policy class includes another, the larger class has at least the capacity of the smaller. Neither says what a system will find. That depends on where the evaluator's errors sit among the behaviors it can reach and on how well search locates them, which is what the finite-output illustration varies: move the scoring defect and each method favors different behavior. The excerpt does not say the ranking of methods reverses, and it does not give the outputs, the scorer or the methods compared.

Why it is useful (my reading). The tempting shortcut is that the substrate that moves the policy least is the safest. The paper's move is to say a bound is not a forecast and to keep the statements separate so nobody reads one as the other. The point is one a practitioner can use without the formalism: a single number for "how much this method changes the model" does not tell you how exposed the method is to a given evaluator's mistakes, because exposure is set by where the mistake is.

The limit. A finite-output illustration is a constructed case. It supports the claim that no ranking follows from distance alone, and it is not evidence about any deployed system or about which substrate is worse in practice. The excerpt has no formulas, so what "distance" is measured from, and whether the bound is tight, is not stated.

Inquiring lines that read this note 51

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What attack surfaces do reasoning traces and chains introduce? Can we reliably detect when models game evaluations? Do honeypot benchmarks validly measure reward hacking better than standard tests? Why do locally safe actions create system-level safety gaps? How do capability benchmark scores systematically misrepresent true model abilities? How do evaluation practices shape which failures stay visible? How do we enforce security boundaries in evaluation environments? What should agent evaluation prioritize to reveal reliable behavior? Can causal models help detect and locate hidden sandbagging in AI? Do backend defenses obscure real attack effectiveness in reported metrics? What determines whether deployed AI systems can actually be stopped in practice? Can single-point security defenses protect multi-agent systems from multi-step attacks? Can local safety checks guarantee system-level behavioral safety?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 85 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

distance alone cannot establish a universal ranking of vulnerability to reward hacking across substrates — in the paper's finite-output illustration the location of the scoring defect changes which behavior each method favors