Does model collapse depend on how we schedule training data?
Does replacing real data with synthetic data each generation cause inevitable model collapse, or is collapse avoidable through different training schedules? This matters because it determines whether training on generated content is fundamentally limited.
Gerstgrasser et al. test whether model collapse is an artifact of how prior studies built their training loops rather than an inherent property of training on synthetic data. Prior work assumed "each model's generated data replaces previous data" — a setup the authors call realistic-sounding but actually inconsistent with practice, since real LLM training sets grow across generations (1.4 trillion tokens for Llama 1, 2 trillion for Llama 2, 15 trillion for Llama 3) rather than swapping old data for new. Pretraining sequences of causal transformers (GPT-2 9M; Llama2 12M/42M/125M) on TinyStories, and resampling synthetic data from each generation's model, they report: "we confirm that replacing the original real data by each generation's synthetic data does indeed tend towards model collapse," while "accumulating the successive generations of synthetic data alongside the original real data avoids model collapse." The same contrast holds for GeoDiff diffusion models on molecular conformations and VAEs on images, "across a range of model sizes, architectures, and hyperparameters."
The mechanism comes from an analytically tractable linear-regression framework (extending Dohmatob et al., 2024a) in which a sequence of linear models is fit to the previous model's outputs. Under replacement, prior theory showed test error grows linearly with the number of fitting iterations. The authors "extend this argument to prove that if data instead accumulate, the test error has a finite upper bound independent of the number of iterations" — collapse is a property of the replacement schedule, not an inevitable consequence of training on model-generated text. Ablations narrow the claim further: growing a synthetic-only dataset to match the accumulated dataset's size still degrades performance, just more slowly, and varying how much training happens in the first iteration has no significant effect. So the result isn't simply "more data helps" — it's that keeping the original real data present, generation after generation, is what bounds the error.
This sits in direct tension with Does training on AI-generated content permanently degrade model quality?, which frames collapse as irreversible tail-loss and already flags — without detailing — "Gerstgrasser et al. 2024" as counter-evidence that the debate isn't settled; this note is that paper. The two findings aren't describing the same experiment: the tail-loss note's irreversibility holds under a replace-data regime (the Shumailov et al. setup); this paper shows that irreversibility is conditional on that regime and disappears once generated data is concatenated with, rather than substituted for, the real corpus. It also complicates How do quality, diversity, and complexity affect synthetic data differently?: where QDC locates generalization effects in properties of the synthetic data itself (diversity, complexity), this paper locates the collapse/no-collapse outcome in the accumulation schedule — whether real data stays in the mix — rather than in any property of the synthetic data generated.
The study is small-scale and deliberately synthetic: kindergarten-level TinyStories text, models up to 125M parameters, and a "maximally pessimistic" assumption that synthetic data is dumped uncontrollably with no filtering or curation. It does not test real web-scale pretraining mixtures, does not address deterministic (temperature-0) generation (left to future work), and only resolves one of four phenomena the authors themselves group under "model collapse" — unbounded test-error blowup — leaving modal collapse, collapse to uniformity, and artifact amplification untouched. The implication the evidence supports is conditional: the curse of recursion is avoidable in principle as long as real human-generated data keeps accumulating alongside synthetic output rather than being crowded out by it — a bound that depends on real data continuing to be collected, not merely having existed historically.
Inquiring lines that read this note 4
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do models learn from self-generated outputs without cascading failures? How do training data quality and composition affect downstream model performance?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does training on AI-generated content permanently degrade model quality?
When generative models train on outputs from previous models, do the resulting models lose rare patterns permanently? The question matters because future training data will inevitably contain synthetic content.
names this paper as unresolved counter-evidence to its irreversibility claim; this note supplies the mechanism and bound
-
How do quality, diversity, and complexity affect synthetic data differently?
When training models on synthetic data, do quality, diversity, and complexity each play distinct roles in how well models generalize? Understanding their separate effects could explain why current optimization strategies fail.
both locate the collapse/generalization outcome in how synthetic data is composed or accumulated, not in synthetic data per se
-
Does verifier filtering actually prevent model collapse long term?
When synthetic training data is filtered through a verifier, does it genuinely solve model collapse or only delay it? The question matters because verifier-guided retraining is a common strategy in modern ML pipelines.
Qualifies A: verifier-filtered retraining only escapes collapse short-term, then converges to the verifier's bias, unlike data accumulation
-
Does synthetic content in search results hide ecosystem decay?
As AI-generated content dominates search rankings, do traditional accuracy metrics mask a silent loss of source diversity and ecosystem health? This matters because hidden fragility could make systems vulnerable to future corruption.
Extends A's model collapse to retrieval: synthetic content homogenizes corpora while high answer accuracy masks the diversity loss
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data
- Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term Convergence
- When AI Reviews Its Own Code: Recursive Self-Training Collapse in Code LLMs
- Scaling Synthetic Data Creation with 1,000,000,000 Personas
- Reasoning-Driven Synthetic Data Generation and Evaluation
- The Curse Of Recursion: Training On Generated Data Makes Models Forget
- Foundation Priors
- CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks
Original note title
model collapse is avoided when synthetic data accumulates alongside real data instead of replacing it