Does code LLM self-review prevent recursive training collapse?
When code models review their own generated outputs across multiple training rounds, can self-scoring or perplexity filters maintain quality, or do they eventually rubber-stamp degraded code? Understanding self-gate failure modes matters for safe recursive training.
Across four code LLMs (SantaCoder, Qwen2.5-Coder, StarCoder2-3B, Code Llama-7B) and five rounds of recursive fine-tuning on their own generated code, the paper compares three review regimes: no review, "Human-gate" review using model-independent signals (compilation, static quality rules), and "AI-self-gate" review using the code LLM's own perplexity or binary self-scoring. The result: "no review collapses fastest, Human-gate filters slow but do not stop collapse, and AI-self-gate filters can look strong early but later lose their filtering effect." In the clearest case, "the binary self-gate enters a rubber-stamp regime where acceptance scores rise while benchmark correctness falls" — on MBPP+, the perplexity filter's pass rate rises from 0.167 at round 1 to 0.235 at round 5 even as the underlying model degrades.
The paper's mechanism is that a Human gate uses an acceptance function rH(x,c) that is fixed and "independent of t and θt" — exogenous to the generator — while an AI self-gate's acceptance score rφt is produced by the same model family being retrained, so it "can drift as the generator changes." They prove AI self-gating "degenerates to ungated self-training under a self-confirming acceptance condition": once the generating model has already drifted, its own perplexity or self-score becomes miscalibrated in the same direction, so the filter stops filtering. A spectral analysis of representation covariance shows the leading variance direction gets amplified round over round relative to the rest, concentrating the output distribution regardless of which gate is used — gates only change the rate, not the eventual "collapse attractor" all four models converge toward.
This sharpens What limits how much models can improve themselves?: that framework treats the verifier as a fixed quantity whose gap with generation sets the ceiling, but here the "verifier" is the generator's own scoring head, so the gap itself erodes over iterations rather than holding constant. It also qualifies Can models reliably improve themselves without external feedback? — external anchoring is the right direction, but this paper shows a weak exogenous anchor (compilation, static rules) is not sufficient either: Human-gate filtering "preserve[s] useful validity signals but cannot stop long-horizon semantic drift." Unlike How quickly do errors compound during model self-training?, which concerns a fully unverified loop, this paper's finding is that even reviewed loops collapse once the reviewer shares the generator's distribution.
The experiments stop at five rounds, Python benchmarks, and models up to 7B parameters (with additional StarCoder trajectories only in an appendix), and the paper states its Human-gate filters are "simplified proxies for full PR review" that "do not fully check semantics, security, or design quality" — so the result does not establish that stronger, more human-like PR review would also fail, only that compile/static checks and self-scoring both do. The paper's own conclusion is narrower and more actionable than a general indictment of self-review: "stable recursive code LLM training requires exogenous verification rather than model-coupled self-review," which licenses caution specifically about AI-code-reviewing-AI-code pipelines, not about human review generally.
Inquiring lines that read this note 1
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do models learn from self-generated outputs without cascading failures?Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
What limits how much models can improve themselves?
Explores whether self-improvement has fundamental boundaries set by how well models can verify versus generate solutions, and what this means across different task types.
shows the gap itself decays when the verifier is coupled to the generator, rather than holding fixed
-
Can models reliably improve themselves without external feedback?
Explores whether self-improvement alone can sustain progress or if structural limits—like the generation-verification gap and diversity collapse—require external anchoring to work reliably.
qualifies the external-anchoring fix: a weak exogenous anchor (compile/static checks) slows but does not stop collapse
-
How quickly do errors compound during model self-training?
When LLMs train on their own outputs without verification, do small mistakes amplify exponentially? This matters because it determines whether unsupervised self-improvement is even feasible.
contrast: that note's loop has no review at all, while this paper shows reviewed loops collapse too once the reviewer is model-coupled
-
Does training on AI-generated content permanently degrade model quality?
When generative models train on outputs from previous models, do the resulting models lose rare patterns permanently? The question matters because future training data will inevitably contain synthetic content.
repository-scale instance of the same recursive-training collapse, specific to code and to review-gate design
-
Why do self-written tests pass when deployment fails?
When agents write and grade their own tests, they can achieve high scores while real performance lags or regresses. This explores the gap between self-reported success and actual deployment behavior.
Evidence for: self-written tests pass despite deployment regression, paralleling self-review's rubber-stamping — only external audits close the gap
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- When AI Reviews Its Own Code: Recursive Self-Training Collapse in Code LLMs
- LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails
- Large Language Models Cannot Self-Correct Reasoning Yet
- A Self-Improving Coding Agent
- Understanding Before Reasoning: Enhancing Chain-of-Thought with Iterative Summarization Pre-Prompting
- Can Large Language Models Really Improve by Self-critiquing Their Own Plans?
- When Hindsight is Not 20/20: Testing Limits on Reflective Thinking in Large Language Models
- Self-Authored Verification Is Unreliable in Heuristic Self-Improving Agents
Original note title
AI self-gate review of generated code degenerates into rubber-stamping while human-gate filters only slow collapse