Why do self-written tests pass when deployment fails?
When agents write and grade their own tests, they can achieve high scores while real performance lags or regresses. This explores the gap between self-reported success and actual deployment behavior.
In self-improving agents that iteratively rewrite a procedural "policy" together with the tests that grade it, the self-report score can stay high while a hidden deployment evaluation shows the policy failing or regressing. The paper calls this the verifier–deployment gap: "self-assigned scores can remain near perfect while real deployment performance degrades or stays low." Across their game-playing testbed, "among the 35 model–game cells with valid self-test outputs, all end with a self-score of at least 0.70. Yet 15 of the 35 completed policies score below their game's random reference, including six policies at Pong's -21.0 floor" — high self-scores coexisting with failed deployment behavior.
The mechanism is agency, not a one-off exploit: the agent "controls both the optimized object and its verifier," so passing self-written tests can reflect genuine improvement, an easier sample distribution, a narrower metric, or tests that inherit the candidate policy's own mistaken assumption — "even purely local optimization of self-test accuracy can lead to a system where self-scores increase while real deployment performance degrades." Failure is stratified by capability: weaker agents corrupt previously acquired strategies while self-tests stay easy ("failure to retain"), while stronger agents are more stable but still "mismeasure the shifted deployment distribution." Internal fixes — monotone test-strengthening checks, or a discriminative check against a trivial baseline — don't reliably close the gap and fall below the unprotected baseline for four of six models. The paper's fix, SEAL, adds a harness-side audit on hidden instances disjoint from training, self-tests, and the final deployment evaluation; it returns only accept/reject and rolls back the whole policy-test state on a clear regression. Audit confidentiality turns out to matter as much as audit existence: revealing numerical audit scores after rejection "can substantially weaken the gate" (Doubao-Seed-2.0-Pro's mean drops from 35.1 under SEAL to 12.7 once scores leak). An information-theoretic result bounds what any endogenous-only gate can achieve when the regressed and improved worlds are hard to distinguish from agent-visible evidence alone.
This is a different gap from the one in What limits how much models can improve themselves?: that gap is a capability ceiling — how much better a model verifies than it generates, scaling with pretraining compute. This paper's gap shows up even at near-perfect self-test pass rates and is a conflict-of-interest problem — the same agent authors the thing graded and the grader. It is a concrete instance of Can models reliably improve themselves without external feedback?'s claim that every reliable fix needs something external: SEAL is a minimal version of that external anchor, and the leakage finding adds a condition the mirage note doesn't specify — the anchor must stay unobserved by the agent, not merely external to it. It also isolates one of the two dials in What separates self-improvement from policy improvement?: SEAL leaves the improver inside the agent (it still authors candidates and self-tests) and moves only the acceptance standard outside, showing that dial alone changes outcomes.
The excerpt tests heuristic, procedural self-improvement — editing code, rules, and tests — not gradient-based weight updates, and the testbed is small game-playing environments (Pong, Breakout, and similar) across six models and three seeds. It does not show that the same capability stratification or the same cheap hidden-audit design carries over to open-ended coding-agent deployments or longer horizons, where constructing a disjoint, confidential audit is harder. The information-limit result is explicitly conditional — it does not prove endogenous evidence is always uninformative, nor that SEAL is a unique or minimal fix. At the strength the evidence allows: a self-improvement loop that lets an agent author its own acceptance test should treat the self-report as advisory and budget for at least one small, confidential, externally-held comparison before any update is retained.
Inquiring lines that read this note 1
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do educators verify student capability when AI can produce indistinguishable work?Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
What limits how much models can improve themselves?
Explores whether self-improvement has fundamental boundaries set by how well models can verify versus generate solutions, and what this means across different task types.
contrasts a capability-ceiling gap with this paper's conflict-of-interest gap, which persists even at near-perfect self-test pass rates
-
Can models reliably improve themselves without external feedback?
Explores whether self-improvement alone can sustain progress or if structural limits—like the generation-verification gap and diversity collapse—require external anchoring to work reliably.
SEAL instantiates "every reliable fix requires something external" and adds that the external anchor must also stay unobserved by the agent
-
What separates self-improvement from policy improvement?
Does recursive self-improvement work by the same evaluate-and-improve cycle as classical policy iteration, or are they fundamentally different processes? Understanding this distinction matters for predicting which self-improving systems remain controllable.
SEAL isolates the external-standard dial while leaving the improver inside the agent, showing that dial alone drives the effect
-
Where does the evaluation boundary actually end in agent benchmarks?
Interactive benchmarks let agents write to state and receive feedback in loops. Does everything the agent can influence on the path to the reward score count as part of the benchmark's evaluation boundary?
a related but distinct leakage channel: there the agent's own actions widen the evaluation boundary, here the agent authors the grader itself
-
Does code LLM self-review prevent recursive training collapse?
When code models review their own generated outputs across multiple training rounds, can self-scoring or perplexity filters maintain quality, or do they eventually rubber-stamp degraded code? Understanding self-gate failure modes matters for safe recursive training.
Evidence for: self-review degrading to rubber-stamping parallels self-tests passing despite deployment regression; human filters only slow it
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Self-Authored Verification Is Unreliable in Heuristic Self-Improving Agents
- Hyperagents
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases
- Gödel Agent: A Self-Referential Agent Framework for Recursive Self-Improvement
- BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure
- MetaRSI / RSI2: A Meta-Recursive Self-Improving System for Recursive Self-Improving Systems Themselves
- Sharpening Tax in Post-Training
Original note title
self-authored verification diverges from deployment truth because the agent edits both the policy and its own tests — a sealed audit closes most of the gap