Optimizing the Score, Losing Sight of the Task: Reward Hacking Across Weights, Selection, and Prompts
A higher evaluation score does not always mean a better language model system. When optimization exploits an evaluator’s mistakes, measured progress can conceal unchanged or deteriorating task performance. This failure can arise through parameter updates, selection among generated outputs, or revisions to persistent prompts. We develop a comparative framework for reward hacking across these three optimization substrates: weights, selection, and text. Building on the Proxy Compression Hypothesis and research on inference-time and in-context reward hacking, we examine how reachable behavior, optimization budgets, and persistent adaptation shape exposure to proxy error. We formalize a distance-dependent upper bound on evaluator disagreement and a capacity ordering for nested policy classes, then show why distance alone cannot establish a universal ranking of vulnerability. An exact finite-output illustration demonstrates how the location of a scoring defect changes the behavior favored by each method. We also map representative defenses across substrates, identifying which mechanisms transfer directly and which offer only functional analogies.
Introduction. A language model system receives a higher score after an update. Has it become better at the task, or better at satisfying the evaluator? That distinction is central to any system that uses measured performance to guide its own improvement. A score makes optimization possible, but it also creates an opportunity: behavior that exploits a weakness in the measurement can be rewarded alongside behavior that solves the task. This opportunity appears in several familiar forms. Parameter training can increasingly favor outputs that a reward model scores highly even as an independent measure of quality declines (Gao et al., 2023). Best-of-n selection can expose a scorer’s blind spots by searching a larger pool of candidates (Khalaf et al., 2025). Persistent prompt optimization can encode a scoring shortcut into instructions that are then reused across future inputs. In one production example, a prompt mutation raised a judge’s rationale-alignment pass rate from 23.1% to 80.0% by adopting its preferred vocabulary, while defect-identification precision remained unchanged (Wahi, 2026).
Discussion / Conclusion. Reward hacking can arise when weights are updated, when outputs are selected, and when persistent text is revised. The shared mechanism is optimization against a signal that fails to represent the task adequately. The substrate matters because it changes which behaviors are reachable, what information persists, and what can be inspected or constrained. A useful comparison therefore needs more than a measure of how far a policy moves. A distance-dependent error envelope gives an upper bound; class inclusion gives a capacity ordering; the geometry of accessible behaviors and the effectiveness of search determine what a particular system actually finds. Keeping these statements separate makes the framework applicable without assuming that every proxy produces the same curve or that one substrate is always safest. For practitioners, the defense correspondence offers the most immediate use.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can single-point security defenses protect multi-agent systems from multi-step attacks?- Should input defenses be validated separately for each channel?
- Can a single security protection work across different system architectures?
- Which reward hacking defenses work across weight updates and output selection?
- Does generalization from named hacks extend to unnamed hacking strategies?
- Is one optimization substrate always safer than another against reward hacking?
- Which reward hacking defenses transfer directly across weights, selection and text?
- What makes a defense mechanism transfer directly rather than just function analogously?
- What properties must defenses preserve to survive substrate differences in persistence and inspectability?
- Can deterministic guardrails designed for judges transfer to weight-based or text-based reward hacking?
- Why does a series of improving scores differ from a single score rise?
- What would a diagnosable evaluation look like compared to a scalar score?
- How should outcomes be scored when comparing applications with different interaction formats?
- Can an average-case validator score hide poor performance on critical tasks?
- When does measured progress on an evaluator conceal actual performance decline?
- How does a ranked default score compete with deliberately optimized outputs?
- Does measured performance gain reflect true task improvement or evaluator exploitation?
- How do hidden partitions in evaluators compare across training and selection substrates?
- How much of weak-to-strong performance gaps reflect presentation rather than capability?