Why does an AI that trains on its own code output get worse with every round, instead of better?
What causes code quality to degrade across multiple rounds of recursive self-training?
This explores why a code model trained on its own output gets worse each round, and what in the training loop causes that drift, as opposed to the cases where self-training actually helps.
This explores why a code model trained on its own output gets worse each round, and what in the loop causes that drift. The most direct answer in the corpus is that the quality filter decays along with the model. In a study of four models over five training rounds, letting the model score its own code looked like it worked at first. But as the model drifted, its judgment drifted with it, and the filter ended up approving worse and worse code Does code LLM self-review prevent recursive training collapse?. Outside checks such as compilers and static rules slowed the decline but couldn't stop it. Code that compiles and passes lint can still be wrong, so the deeper loss of meaning got through.
Why would a model's self-review fail like this? One reason is that models tend to trust answers they generated themselves. An answer the model found likely to produce also looks likely to be correct when it evaluates it Why do models trust their own generated answers?. In a recursive loop, that bias compounds: the model produces code, approves it, trains on it, and becomes even more convinced that this style of code is right. A related effect shows up inside a single task. When a model's earlier mistakes fill its context, its later error rate rises sharply, and bigger models don't escape this Do models fail worse when their own errors fill the context?. Recursive self-training looks like a slower version of the same problem, with the errors stored in the weights instead of the context.
A second cause is narrowing. Each round of keeping only the outputs the model rates highly cuts off the unusual but valid solutions, and diversity shrinks. Work on critique models describes this as "tail narrowing" across self-training rounds. Step-level critique during training helps keep a wider range of solutions alive Do critique models improve diversity during training itself?. In reinforcement learning, the end of that path has been observed directly: when a model's attempts at a prompt all score about the same, the learning signal weakens and the model settles into generic templates that ignore the input Why do language models collapse into generic templates?. A related finding on training models to correct themselves shows that recycled traces don't match the model's actual mistakes, and the model collapses into a single way of correcting Why does self-correction training on offline data fail?.
The counterexamples are the most useful part. Self-training doesn't always degrade. Transformers trained repeatedly on their own *verified-correct* arithmetic went from 10-digit to 100-digit addition, improving faster each round with no plateau Can transformers improve exponentially by learning from their own correct solutions?. The Darwin Gödel Machine kept improving its own coding agent by testing each variant against real benchmarks instead of the model's opinion of itself Can AI systems improve themselves through trial and error?. The difference is the filter: an independent check that doesn't drift with the model, versus self-judgment that does. Another approach filters successful code runs strictly for quality but keeps a varied set of failures as negative examples Why do correct code trajectories teach models to tolerate errors?. That keeps the variety that pure self-selection destroys.
The broader point is that what fails in a self-improvement loop is often the judge, not the generator. Systems that hold up add brakes against drift. One example limits how much an agent can rewrite its own instructions each round, checks every change on held-out tests, and keeps rejected edits on file as examples of what not to do Does constraining edits make skill learning more stable?. Even the form of the checking signal matters: STOP found that its self-improvement loop gained more when the verifier was source code than when it was plain English Can language models improve their own scaffolding without weight updates?. If you want to understand recursive collapse, look first at who grades the work.
Sources 11 notes
Across four models and five training rounds, self-scoring filters initially appear effective but lose filtering power as the model drifts, eventually rubber-stamping worse code. Compile and static-rule checks slow degradation but fail to prevent long-horizon semantic collapse.
LLMs exhibit structural bias toward validating their own outputs because high-probability generated answers feel more correct during evaluation. Comparing answers against broader alternatives breaks this self-agreement loop.
Error accumulation in context causes non-linear performance degradation in long-horizon tasks. Model scaling does not fix this; only test-time compute through thinking models reduces the effect by preventing error-contaminated context from biasing reasoning.
Step-level critique in the training loop counteracts tail narrowing and maintains solution diversity across self-training iterations. This training-time benefit—preventing premature convergence—is more fundamental than test-time accuracy gains.
When within-prompt reward variance is low, task gradients weaken and regularization dominates, pushing policies toward generic outputs. SNR-Aware Filtering—selecting high-variance prompts before updates—recovers performance across tasks and scales.
Show all 11 sources
SFT on offline correction traces fails because training errors don't match test errors and models collapse into single correction modes. Multi-turn online RL under the model's own error distribution successfully trains self-correction by letting models practice correcting their actual mistakes.
Standard transformers generalize from 10-digit to 100-digit addition by repeatedly generating solutions, filtering for correctness, and retraining—showing exponential (not linear) out-of-distribution improvement across rounds without saturation.
DGM replaces formal proofs with empirical benchmarking and maintains an evolutionary archive of agent variants, achieving 2.5× improvement on SWE-bench and 2.2× on Polyglot by discovering capabilities like better code editing and context management.
GRPO-RoC filters positive trajectories for quality while preserving diverse failures as negative signal, allowing a 14B model to reach frontier math performance in 510 RL steps, surpassing much larger models with cleaner reasoning.
SkillOpt's ablations show that adding a textual learning-rate budget, held-out validation gate, and rejected-edit buffer (retaining failed edits as negative feedback) produces more stable and generalizable skill improvement than allowing agents to freely rewrite their own instructions.
STOP demonstrates that an LM can iteratively refine the improver program wrapped around it, achieving measurably better downstream performance without any weight changes. The form of the verifying signal—source code versus plain English—significantly shapes how much improvement the loop can extract.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Large Language Models Cannot Self-Correct Reasoning Yet
- Self-Authored Verification Is Unreliable in Heuristic Self-Improving Agents
- Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models
- Hyperagents
- When AI Reviews Its Own Code: Recursive Self-Training Collapse in Code LLMs
- Training Language Models to Self-Correct via Reinforcement Learning
- Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges
- Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback