When AI Reviews Its Own Code: Recursive Self-Training Collapse in Code LLMs
Recursive self-training can degrade neural generative models when generated data is reused without fresh human data or external quality control. We study this risk in code LLMs, where AI-generated code can enter real repositories, later become training data, and create a repository-scale self-training loop. While software development traditionally interrupts this loop through pull-request review, tests, compilation, and human approval, AI coding tools now produce code faster than humans can review it, and code review itself is increasingly automated by AI systems. We therefore compare three recursive fine-tuning regimes: no review, Humangate review using model-independent filters such as compilation and static quality checks, and AI-self-gate review using the code LLM’s own signals such as perplexity and binary self-scoring. Across multiple code LLMs and benchmarks, no review collapses fastest, Human-gate filters slow but do not stop collapse, and AI-self-gate filters can look strong early but later lose their filtering effect. In the clearest case, the binary self-gate enters a rubber-stamp regime where acceptance scores rise while benchmark correctness falls. We explain this behavior by formulating review as gated distributional reweighting, proving that AI self-gating degenerates to ungated self-training under a self-confirming acceptance condition, and giving a spectral analysis of representation-level covariance concentration under recursive retraining. These results suggest that stable recursive code LLM training requires exogenous verification rather than model-coupled self-review. Codes are available at https://github.com/Hik289/code-retraining.git.
Introduction. Recursive self-training has been repeatedly shown to be harmful for neural generative models when the loop is not anchored by fresh human data or an external quality signal. In image generation, self-consuming training can reduce distributional variance and drive models toward low-diversity outputs, a failure mode described as model amplification disorder Alemohammad et al. [2023]. In general generative modeling, training on recursively generated data can cause model collapse: distributional errors accumulate, low-probability modes are lost, and the learned distribution drifts away from the original data distribution Shumailov et al. [2024], Dohmatob et al. [2024]. Similar effects have been reported for language models, where synthetic text can reduce diversity, amplify earlier mistakes, and degrade long-horizon quality unless real data or strong external filtering is preserved in the training loop Briesch et al. [2023], Seddik et al. [2024], Suresh et al. [2025]. These results give a general warning: self-training is not automatically self-improvement. When a model is trained on data produced by earlier versions of the same kind of model, the loop can become self-reinforcing rather than corrective.
Code large language models (LLMs) Ben Allal et al. [2023], Li et al. [2023], Rozière et al. [2023], Hui et al. [2024] are likely to face the same risk in a concrete and testable form. A code model can generate programs, add them back into the training corpus, and fine-tune the next model on this synthetic code. If these programs were always correct, diverse, and well reviewed, the loop could cheaply scale code data. In practice, generated code often contains bugs, incomplete logic, repeated templates, missing edge cases, and narrow stylistic patterns. Recursive training can therefore copy these errors into later models, making them better at imitating their own code style while worse at solving programming tasks. This is the recursive self-training trap for code: more data is created, but it can carry the model’s own functional errors into future training.
This issue is no longer only an offline training concern. AI-generated code is rapidly entering real repositories through tools such as GitHub Copilot GitHub [2026a], Cursor Anysphere [2026], Amazon CodeWhisperer Amazon Web Services [2026], and Devin Cognition AI [2026]. Studies report large productivity gains from Copilot Peng et al. [2023], and enterprise reports show that Copilot-suggested code is frequently accepted, committed, and merged into pull requests GitHub [2024]. Future code LLMs may therefore be trained on repositories already containing AI-generated or AI-assisted code. If such code is collected without review, the loop becomes recursive self-training at repository scale: AI writes code, the code enters GitHub, GitHub enters the next training corpus, and the next model learns from earlier model outputs.
In real software development, code is usually filtered before entering a repository. A developer opens a pull request (PR), and the change may pass through compilation, tests, static analysis, and human review before merging. Modern code review is a lightweight but consequential qualitycontrol process: it catches defects, transfers project knowledge, and affects downstream software quality Bacchelli and Bird [2013], Bosu et al. [2015], McIntosh et al. [2016], Sadowski et al. [2018]. This PR process acts as an external quality gate: the author does not define the acceptance rule, and the rule does not become weaker when the author writes worse code. Failed builds, failing tests, or poor design can still block the PR. We call this model-independent acceptance mechanism a Human gate. The term does not require every decision to be manual; it means that the verifier is exogenous to the generator and preserves a fixed quality standard.
The AI coding era weakens this assumption. AI tools can generate PRs faster than humans can review them, while AI code review is becoming a practical engineering workflow. Recent studies examine LLM-assisted review with AI co-reviewers and interactive review assistants Aðalsteinsson et al. [2025]; large-scale GitHub Actions analyses show AI review tools automatically commenting on real PRs Sun et al. [2025]; and GitHub Copilot code review is now integrated into the PR workflow GitHub [2026b]. At the same time, LLM-as-judge systems are known to be sensitive to evaluator design and can favor model-like outputs Zheng et al. [2023], Panickssery et al. [2024]. Future repositories may therefore contain both AI-generated and AI-reviewed code. In this setting, the verifier may no longer be a human-quality external process, but another AI model, possibly close to the generator. If AI writes the code and AI reviews the code, the acceptance signal can become coupled to the same distribution being retrained.
Method. Figure 1 illustrates the two recursive training pipelines considered in this paper: ungated self-training, which reuses all generated code, and gated self-training, which filters generated code before retraining. Let x ∈X denote a prompt or programming task, and let c ∈C denote a candidate code solution.
No verification (ungated recursion). In ungated recursive self-training, every generated sample is directly reused as training data. The synthetic data distribution at iteration t is Verification (gated recursion). In gated recursive self-training, generated code is accepted only after passing a review or verification signal. Let E(x, c) ∈{0, 1} denote the acceptance event, and define Here, r(x, c) can represent a compile check, execution test, static quality filter, human review score, or AI-based review score. Conditioning the generator on acceptance gives the accepted distribution Thus, gated recursion differs from ungated recursion only through the sampling measure: before training the next model, generated samples are reweighted by the acceptance signal r(x, c).
We next describe a representation-level view of recursive self-training collapse. Let Dt denote the training distribution at iteration t over prompt–code pairs (x, c) ∈X × C, where Dt may be ungated or gated. Fix a representation map Here, high similarity means the dominant representation subspace is stable, while large distance indicates that recursive training has shifted the main directions of variation.
Thus, if s1 ≫s2, the component aligned with the leading singular direction dominates after repeated retraining. Whenever Σ0 has nonzero energy in this direction, the ratio λ1(Σt)/λ2(Σt) is expected to increase with t because the leading component is amplified roughly by s2t 1 while the next component is amplified roughly by s2t 2 . This gives a representation-space explanation for collapse: recursive selftraining can progressively concentrate variance into a low-dimensional subspace, reducing diversity in the generated code distribution.
We now specialize the generic acceptance function r(x, c) from Equation (2.3). The key is whether the acceptance rule is fixed outside the generator or coupled to the generator during recursive training.
Human gate: exogenous acceptance. A Human gate uses a fixed acceptance function rH(x, c) ∈ [0, 1], rH is independent of t and θt. It induces the accepted conditional distribution The corresponding training distribution is mH t (x, c) = pX(x)qH θt(c | x). This covers modelindependent review signals such as compilation, execution, quality checks, or human PR review.
AI self-gate: endogenous acceptance. An AI self-gate uses a learned reviewer with parameters φt, producing rφt(x, c) ∈[0, 1]. It induces The corresponding training distribution is mA t (x, c) = pX(x)qA θt,φt(c | x). In our experiments, this case corresponds to perplexity filtering and binary self-scoring, where the code LLM evaluates its own generated samples. Unlike rH, the score rφt can drift as the generator changes.
Taxonomy summary. Table 1 classifies the six filtering strategies by whether the acceptance rule is exogenous or endogenous. Vanilla has no gate and accepts all generated code. Compile is a Human gate because it directly runs Python compilation and accepts code only when compilation succeeds. Quality is also a Human gate because it uses fixed code-quality rules, described in Section D, such as repetition rate and length. Compile+Quality combines these two model-independent rules, so it is also exogenous. Perplexity and Binary Classifier are AI self-gates: both use the code LLM itself to score its own generated code. Perplexity filters by the model’s own likelihood, while Binary Classifier uses the model’s own logit-difference score, defined in Section E. Since the generator also provides the filtering signal, these two gates are coupled to θt.
Coupled dynamics under AI self-gating. We now formalize the failure mode where the verifier is coupled to the generator. This captures the AI self-gate setting: the code LLM generates candidate code, scores its own outputs, and then trains on the samples it accepted.
Discussion. Perplexity self-gating loses calibration. Among all AI self-gate methods, perplexity filtering produces a more stable trajectory than binary classification on MBPP+. The filter pass rate for PPL slowly increases from 0.167 at R1 to 0.235 at R5 (Table 2), indicating gradual degradation of the self-calibration signal (Figure 7). This aligns with Theorem 2.3: once the generating model collapses, its perplexity signal becomes miscalibrated.
MBPP+ degrades more severely than HumanEval+ across all models and filter strategies. Figure 8 shows MBPP+ trajectories for all models under all five filters. The most dramatic collapse is observed for SantaCoder under Vanilla self-training, where MBPP+ drops from the baseline 0.294 to 0.008 at Round 4 (−97.3%). Qwen2.5-Coder also collapses severely under Vanilla: MBPP+ from 0.582 to 0.082 at Round 5 (−85.9%). Code Llama-7B is notably more robust to MBPP+ collapse, with MBPP+ under Vanilla remaining at 0.214 at Round 5, reflecting its stronger pretraining foundation.
Effect of filtering strategies. The benefit of a gate depends on which failure mode the benchmark rewards. Compile filtering improves Qwen2.5-Coder at R5 relative to Vanilla on both HumanEval+ (0.098 vs. 0.043) and MBPP+ (0.127 vs. 0.082), and it is the strongest Code Llama-7B strategy on HumanEval+ (0.171). For StarCoder2-3B, however, compile and quality filters lag behind Vanilla at R5 on HumanEval+ and MBPP+, showing that syntactic validity alone cannot guarantee semantic preservation. Binary filtering achieves the highest R5 HumanEval+ for three of four models, which means self-scoring can still rank some useful samples in early and medium horizons. Its MBPP+ behavior is weaker and less stable, matching our central story: AI self-gates may improve one visible metric while losing broader semantic coverage. Perplexity (PPL) filtering shows a slow increase in pass rate from approximately 0.17 to 0.24 (averaged across models; Figure 7), consistent with the endogenous degeneracy result of Theorem 2.3.
Collapse attractor. Despite large differences in architecture size, pretraining data, and starting pass@1, all four models show qualitatively similar long-horizon degradation patterns. SantaCoder (0.171 HumanEval+ baseline) and Qwen2.5-Coder (0.372 baseline) both reach approximately 0.04– 0.10 HumanEval+ pass@1 under Vanilla by Round 5. StarCoder2-3B and Code Llama-7B maintain somewhat higher residual performance (0.10–0.14), likely due to stronger capacity and pretraining data quality. This convergence to a common degradation floor is consistent with the theoretical prediction of a collapse attractor.
Key takeaway. No filtering strategy evaluated here—Compile, Perplexity, Quality, or Binary— prevents collapse across all model families. Binary filtering achieves the best R5 HumanEval+ for three models, but this comes with erratic MBPP+ behavior. Human gates are more interpretable and preserve executable validity, yet simple compile/static filters are still too weak to guarantee semantic
Conclusion. We studied recursive self-training collapse in code LLMs under no review, Human-gate review, and AI-self-gate review. Across models and benchmarks, no review produces the sharpest degradation, Human gates preserve useful validity signals but cannot stop long-horizon semantic drift, and AI self-gates can look strong on early HumanEval-style metrics while their acceptance signal drifts with the generator. Our theory explains this through gated distributional reweighting: exogenous gates preserve an external quality signal, while endogenous gates can become self-confirming. Stable recursive code LLM training therefore needs verification that remains outside the model’s own preference distribution.
Limitations. Our experiments focus on Python benchmarks and models up to 7B parameters in the full five-strategy sweep, with additional StarCoder trajectories in the appendix. Human-gate filters are simplified proxies for full PR review and do not fully check semantics, security, or design quality. Longer horizons, larger models, multilingual code, stronger execution-based gates, and periodic humanverified data remain important future directions.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How do models learn from self-generated outputs without cascading failures?- What causes code quality to degrade across multiple rounds of recursive self-training?
- What happens when models train on AI-generated content recursively?
- What failure modes emerge when model-generated content trains on itself iteratively?
- Can models learn to generate their own training examples effectively?
- What causes irreversible model collapse when training on model-generated content?
- What makes deliberate practice on your own errors more effective than copying others?
- Why do error avalanches accelerate in self-training loops without verification?
- Can self-consistency checks fully prevent error avalanching in self-training loops?
- What reliable traces do generative processes actually leave in finished text?
- Can fabrication of content serve productive purposes in prediction?
- What training data contamination rates threaten model safety most practically?
- How does diversity loss in synthetic data mirror tail distribution disappearance?
- Why do weaker models generate better training data than stronger models?
- How does the ratio of synthetic to real training data affect model collapse?
- Can synthetic data generation work without seed examples?