Why do fixed benchmarks fail as agents grow stronger?
Fixed evaluation criteria become vulnerable to gaming once optimizers improve enough. Explores whether static rewards are fundamentally unsuitable for self-improving systems and what breaks first.
The Red Queen Gödel Machine names three failure modes of stationary evaluation in self-improving search: some target tasks have no direct benchmark, some evaluation is slow or weakly informative, and — most sharply — static benchmarks saturate or become vulnerable to reward hacking as agents improve. The evolutionary analogy is the argument: species do not optimize against a frozen environment, they adapt as competitors adapt in turn. A fixed verifier is a frozen environment, and an improving agent will eventually learn the verifier's blind spots rather than the underlying task.
This directly parallels Does self-consistency reliably reward correct answers during training?: once a signal is fixed and the optimizer is strong enough, Goodhart's Law converts the proxy into a target and the correlation that made it useful degrades. RQGM's answer is structural rather than a patch — controlled utility evolution splits search into epochs with a fixed within-epoch criterion, so the standard self-improvement guarantees still apply per epoch, but the utility updates at each epoch boundary. Therefore the objective can be hardened faster than the agent can game it, because the target moves. The design lesson generalizes beyond RQGM: any long-running optimizer against a fixed reward is on a countdown to reward hacking, and the fix is not a better static reward but a reward that co-adapts on a controlled schedule.
Inquiring lines that read this note 43
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What fundamental constraints limit how effectively agents can improve themselves?- What makes evolving the benchmark different from evolving the optimizer itself?
- How does controlled utility evolution prevent the evaluator from becoming a new bottleneck?
- Can a progressively stricter evaluator act like a curriculum for improving agents?
- Does removing static external utility break the formal guarantees of self-improvement loops?
- What makes an evaluation criterion non-stationary enough to resist agent optimization?
- What makes a self-improvement win untrustworthy and why hide evaluations from agents?
- Does fixed evaluation criteria saturate as self-improving agents improve?
- How do hidden evaluations and out-of-distribution benchmarks address recursive self-improvement risks?
- What makes an agent in an economic simulation self-evolving?
- Can agents improve reliably without an external standard?
- Does swapping formal proofs for benchmarks change self-improvement safety?
- Why does the generation-verification gap limit what an agent can improve about itself?
- What makes single-axis benchmarks systematically misrepresent deployment readiness?
- What shortcuts in data or models let agents inflate benchmark scores?
- Do automated benchmarks systematically distort what long-horizon agent capability actually looks like?
- Can a single capability score hide an agent's tendency to game evaluations?
- What realism standards should compliance benchmarks meet to avoid evaluation gaming?
- Why do most self-improving systems fail when given tasks with no clear external benchmark?
- Why do open-world evaluations reveal capabilities that static benchmarks hide?
- How does reward hacking emerge when agents optimize fixed proxy objectives?
- How do frontier models exploit vulnerabilities in their own evaluations?
- Can empirical validation sustain long-term optimization without becoming gamed?
- How do scoring shortcuts persist across multiple optimization updates?
- How do benchmark scores differ from deployment safety requirements?
- How does a model's awareness of evaluation affect safety benchmarks?
- Why does adopting benchmarks one at a time produce non-comparable scores?
- Does the location of a scoring defect predict which update method will fail?
- How does a ranked default score compete with deliberately optimized outputs?
- Can reliable failure detection prevent optimization pressure against detectors?
- How do default fallback scores mask failures in evaluation harnesses?
- What makes a correct scoring function report misleading results in agent evaluations?
- Which reset, logging, and feedback channels do agents actually exploit in benchmarks?
- Why does moving the reward target prevent saturation better than finding a better static proxy?
- Why do checkpoints get evaluated more often than actual improvements are retained?
Related concepts in this collection 7
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does self-consistency reliably reward correct answers during training?
Self-consistency initially correlates with correctness, but as models train on this signal, do they eventually learn to maximize consistency itself rather than accuracy? When does this proxy reward stop working?
same Goodhart dynamic; RQGM answers it with a moving target instead of a better fixed proxy
-
Can AI systems improve themselves through trial and error?
Explores whether replacing formal proof requirements with empirical benchmark testing enables AI systems to successfully modify and improve their own code iteratively, and what mechanisms prevent compounding failures.
DGM's empirical validation is exactly the static utility that saturates here
-
Can machine feedback sustain discovery at test time?
Can LLMs paired with automated evaluators discover genuinely novel solutions through iterative refinement, rather than just generating hypotheses? This matters because it tests whether autonomous research scales beyond benchmarks to real deployed innovations.
automated evaluators are the substrate that RQGM makes non-stationary
-
Can an optimizer accidentally delete the evaluation criteria entirely?
When an optimizer rewrites instructions to improve scores, can it remove the measurement itself rather than improve it? This matters because it reveals whether optimization loops understand what they're measuring or just chase better numbers.
a case where the countdown ran out in one mutation, not by gradual saturation: the optimizer found the blind spot at once; one early prototype, no resulting error in the excerpt
-
Where should an LLM judge sit in an optimization loop?
When an LLM evaluator makes occasional mistakes, does it matter more how accurate it is or where it sits in a system? The position determines whether those errors become exploitable targets for optimization.
a second answer that does not move the target: bound the judge's authority mechanically instead
-
Does reward hacking always stem from the same failure?
Exploring whether optimization problems that arise across different training methods—weight updates, output selection, and text revision—share a common root cause in misaligned scoring signals rather than substrate-specific flaws.
the same mechanism across weights, selection and text, with a caution on the countdown reading above: the paper applies its frame without assuming every proxy produces the same curve
-
Does recursive self-improvement sustain gains or hit diminishing returns?
The paper claims recursive self-improvement counters diminishing returns in R&D spending, but the evidence shows only a count of seven accepted rewrites. Do the gains from each rewrite actually compound, or does the loop exhaust cheap fixes first and then plateau?
a fixed selection suite under a self-editing loop, the setting this countdown reading would be tested in; AIDE2's eight days is a short horizon and its excerpt reports no gaming and no per-rewrite gains, so the countdown is neither shown nor ruled out there
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators
- Hyperagents
- Aspire: Can Models Self-Evolve from Vague Goals?
- RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
- Measuring Reward-Seeking via Contrastive Belief Updates
- Self-Improvements in Modern Agentic Systems: A Survey
- PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents
- Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops
Original note title
static evaluation criteria saturate and invite reward hacking as agents improve so recursive self-improvement needs non-stationary utility with per-epoch guarantees