How can we make reward-hacking visible in agent evaluation?
Typical benchmark scores collapse multiple factors into a single number, hiding whether agents are genuinely solving tasks or exploiting reward signals. Can separating evaluation components expose these failure modes?
AgentCompass argues the agent-evaluation landscape suffers an infrastructural deficit: benchmarks operate as isolated suites, each forcing researchers to re-configure heterogeneous environments, data formats, and scoring protocols. This redundant engineering hurts efficiency but — more importantly — compromises reproducibility, because inconsistent baseline implementations mean two labs reporting the "same" benchmark aren't running the same evaluation. The fix is architectural: decouple the pipeline into independent Benchmark, Harness, and Environment components so configurations can be swapped without reimplementing execution logic, backed by a fault-tolerant asynchronous runtime.
The deeper payoff is diagnostic. Because the harness is a separated component with comprehensive trajectory analysis, the infrastructure can go beyond scalar scores to surface how an agent behaved — including nuanced failure modes like reward-hacking that a final-accuracy number conceals entirely. This mirrors the conceptual separation in What are the three distinct layers of agent code?: treating the harness as a first-class object rather than glue code is what lets you attribute behavior. A scalar score collapses model capability, harness scaffolding, and environment quirks into one indistinguishable number, which is precisely why reward-hacking stays invisible under it. Separating the components is therefore not just tidy engineering — it is the precondition for How should we measure agent system performance beyond task success?, because you cannot analyze a trajectory you never isolated.
Inquiring lines that read this note 91
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do pretraining biases affect reward signal effectiveness in RLVR?- Can system-level engineering fixes replace hand-designed reward heuristics entirely?
- What makes a reward evolution schedule fast enough to outpace exploitation?
- How should harness scaffolding be treated as a first-class object?
- What persistent failures remain unsolved despite harness evolution efforts?
- What feedback signals matter most during harness evolution search?
- How do agentic systems hide harness failures from benchmarks?
- How does reward hacking emerge when agents optimize fixed proxy objectives?
- Why does length exploitation emerge as a reward hacking failure in distillation?
- What are AIDE2's hidden evaluations hidden from, and what makes a win untrustworthy?
- Can infrastructure monitoring catch reward hacking that scores cannot detect?
- Can hidden test sets reveal reward hacking that single public scores conceal?
- How is a reward hack defined and labeled across different benchmark studies?
- What determines the ground truth when detecting reward hacking in model evaluations?
- What specific reward-hacking shortcuts did frontier models find in Terminal Bench?
- How does reward hacking differ from errors in the scoring function itself?
- What rates of reward hacking occur in frontier language model benchmarks?
- Do agents that recognize their own reward hacking say so in what they hand back?
- How does optimization pressure against monitors change the visibility of reward hacking?
- Why do agents show awareness of reward hacking but continue doing it?
- Do agents aware of their own reward hacking disclose it in outputs?
- Do agents recognize their own reward hacking before submitting their answers?
- Can co-evolving evaluators alongside actors prevent reward hacking?
- Do agents frame reward hacks as valid strategies rather than flaws?
- What ground truth labels should define reward hacking in automated detection?
- Why does decoupling evaluation components make reward-hacking diagnosable?
- Do models reward hack at high rates on unmodified benchmarks?
- Can belief checks detect whether models will resist reward hacking?
- How often do deployed models exploit evaluation environments to hack their scores?
- Do agents disclose reward hacking in the outputs they return?
- Is one optimization substrate always safer than another against reward hacking?
- Does reward hacking always make capability appear stronger than it is?
- What distinguishes reward hacking from genuine targeting of the grading process?
- How does chain-of-thought monitoring fail when agents hide reward hacking?
- Does reward hacking cause evaluations to overstate model capabilities?
- What makes a win untrustworthy and how does AIDE2 avoid spurious optimization?
- How can reward metrics distinguish novel methods from shortcuts aimed at the evaluator?
- What evidence should benchmark operators attach to completion claims?
- When does an agent's action earlier in the loop change what a scorer reads later?
- Why does held-out evaluation matter for detecting agent overfitting?
- Why do outcome-level metrics fail to reveal contained attacks in multi-agent pipelines?
- How much of an agent's behavior actually escapes human review in practice?
- Does endpoint-only scoring hide meaningful progress like the Judgment Bypass Rate found?
- What makes a win untrustworthy in hidden evaluation environments?
- How can reviewers be matched on effort when monitoring reveals different amounts of behavior?
- What makes a correct scoring function report misleading results in agent evaluations?
- Which reset, logging, and feedback channels do agents actually exploit in benchmarks?
- Can per-stage results reveal which demand causes agent failure in exploitation?
- How do agent actions change state that reward procedures later read?
- Can an agent change reward-path state through actions during evaluation?
- Should feedback channels be excluded from the reward path in agent evaluations?
- How do evaluation hacks differ from genuine sandbox escapes?
- Why does decoupling evaluation into components make hacking more diagnosable?
- What shortcuts in data or models let agents inflate benchmark scores?
- What agent evaluation dimensions beyond task success does a single number hide?
- Can a single capability score hide an agent's tendency to game evaluations?
- Can a single leaderboard score capture multi-dimensional differences in agent performance?
- What realism standards should compliance benchmarks meet to avoid evaluation gaming?
- How does dense task grading compare to honeypot detection for evaluating real capability?
- Do unmodified benchmarks expose similar reward hacking vectors without planted shortcuts?
- Do planted test cases reliably detect agent hacking behavior?
- Can intermediate primitives be scored separately in exploitation benchmarks?
- How do planted detectable hacks compare to human inspection of agent traces?
- What methods could find unplanted hacks that benchmark designers missed?
- How visible is the optional shortcut to the agent during evaluation?
- Why do honeypot tasks reveal reward hacking better than standard benchmarks?
- How can we detect whether an agent recognized its own reward hacking?
- How do planted hacks in benchmarks help measure real-world reward hacking?
- Can a low exploitation benchmark score indicate refusal rather than inability?
- How does a single score mix exploitation ability with task capability?
- How should benchmarks balance verifiability against outcome resolution?
- Why do hidden test partitions matter more than open evaluation sets?
- When does measured progress on an evaluator conceal actual performance decline?
- How can a second performance metric reveal shortcuts that a single metric would hide?
- Does measured performance gain reflect true task improvement or evaluator exploitation?
- How do hidden partitions in evaluators compare across training and selection substrates?
- What makes a public-versus-hidden test score gap a useful hack indicator?
Related concepts in this collection 6
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
What are the three distinct layers of agent code?
Does separating agent code into model capabilities, system harness, and agent-created artifacts help explain why agentic systems fail and where to intervene for improvement?
grounds: same insistence on treating the harness as a separable first-class component
-
How should we measure agent system performance beyond task success?
Current evaluation metrics collapse agent behavior into a single success score, hiding critical information about how agents operate. What dimensions—trajectory quality, memory use, context efficiency, verification cost—should benchmarks actually measure?
extends: component decoupling is the enabling condition for trajectory-level evaluation
-
Does a single benchmark score actually predict agent readiness?
Single-axis benchmarks rank models by one capability—like task success—but ignore privacy, duration, operating mode, and ecosystem fit. Can one number really capture what matters for deployment?
relates: both argue scalar/single-axis evaluation systematically hides what matters
-
Which security protections actually slow down agent exploits?
ExploitGym varies defenses across 898 real-world instances to isolate how each protection affects agent performance. Understanding which defenses matter most to agents versus humans is critical for defenders.
extends: a separated environment component can also be varied on purpose, so decoupling buys causal attribution to environment properties as well as diagnosis of agent behavior (excerpt-only, no protection effects reported)
-
Can planted honeypots reliably catch reward hacking automatically?
Post hoc inspection by humans or judges can miss reward hacks. The question explores whether embedding detectable hacks into tasks—so agents trigger them if they cheat—offers a more reliable detection method than judging their behavior traces afterward.
contrasts: a second route to diagnosing reward hacking that reads a planted event instead of the trajectory, from a paper that calls LLM-judge trace inspection the underdeveloped part of current practice; a tension is filed in ops/tensions/ asking whether the detector, not the trace access, is what fails (excerpt-only, no detection rates for either route)
-
Does a hacked benchmark score hide what the model actually did?
When models exploit evaluation procedures rather than solving the intended task, their scores conflate two separate abilities—the capability being tested and the ability to game the system. This makes benchmark scores unreliable guides to actual model performance.
exemplifies: the collapse a scalar performs, capability and exploit ability in one number, stated by a paper that answers with a readout on the model's activations where this note answers with a separated harness and trajectory analysis (relayed premise; neither excerpt compares the two)
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations
- BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure
- AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
- Hack-Verifiable Terminal Bench: Evaluating Reward Hacking in Terminal Tasks
- BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks
- Optimizing the Score, Losing Sight of the Task: Reward Hacking Across Weights, Selection, and Prompts
- Natural Emergent Misalignment From Reward Hacking In Production RL
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
Original note title
decoupling agent evaluation into benchmark harness and environment components is what makes reward-hacking diagnosable rather than hidden inside a scalar score