INQUIRING LINE

Does AI training on AI-made content break models because the data is bad, or because it crowds out real human data?

Is model collapse a property of data replacement or synthetic data itself?

This explores whether models degrade when trained on AI-generated data because the synthetic data is inherently harmful, or because of how it gets used, specifically when it pushes the original human-written data out of the training mix.


This explores whether model collapse comes from synthetic data being harmful in itself, or from the training practice of letting synthetic data push real data out. The clearest answer in the corpus points to replacement. When each new model generation is trained only on the previous generation's outputs, test error keeps growing without limit. When synthetic data is added on top of the original real data and the real data is kept, error stays bounded. This holds in a mathematical proof and in experiments on language models, image diffusion models and VAEs Does model collapse depend on how we schedule training data?. So the disaster scenario depends on the schedule. Synthetic data alone does not doom a model.

That doesn't make synthetic data neutral, though. One useful way to think about it: an LLM's output is not a fresh observation of the world. It is a sample from the model's own learned beliefs, shaped by whoever wrote the prompt Should we treat LLM outputs as real empirical data?. Training on it partly means training on your own assumptions. That's why replacement is so damaging. Once real data leaves the loop, nothing outside the model is left to correct its drift.

A common fix is to filter synthetic data through a verifier and keep only the outputs it approves. This helps at first, because quality goes up and noise goes down. Over time, though, the model gets pulled toward whatever the verifier believes, including its blind spots, and the early gains level off and then decline Does verifier filtering actually prevent model collapse long term?. Filtering doesn't remove the problem. It swaps the generator's biases for the verifier's.

The less obvious finding is that collapse-like narrowing doesn't need synthetic data at all. Chat models often give bland, repetitive answers, a problem called mode collapse. One line of research traces it to human preference data: annotators systematically favor familiar-sounding text, and training bakes that bias in Where does mode collapse in language models really come from?. In reinforcement learning, models can also fall back on generic, one-size-fits-all reasoning templates when the training signal stops telling good answers apart from bad ones Why do language models collapse into generic templates?. In both cases the narrowing comes from a feedback loop that stops supplying new, distinguishing information. The source of the data doesn't matter.

So the corpus suggests that 'synthetic versus real' is the wrong question. A better one is whether anything independent of the model still enters the training loop. Keeping real data, building diversity deliberately, and treating generated text as a weighted prior rather than as evidence all address that. How well synthetic data works also varies by domain, model and scale, so there's no universal safe recipe What makes synthetic data work across different domains and models?.


Sources 6 notes

Does model collapse depend on how we schedule training data?

Replacing real data with synthetic data causes unbounded test error growth, but accumulating synthetic data alongside the original real corpus keeps error bounded across model architectures and sizes. The mechanism is proven analytically in a linear-regression framework and confirmed empirically on language models, diffusion models, and VAEs.

Should we treat LLM outputs as real empirical data?

Foundation Priors framework shows that LLM-generated text reflects the model's learned patterns and user's prompt choices, not ground truth. Such outputs should only influence inference through explicitly parameterized trust weights, not be treated as equivalent to real evidence.

Does verifier filtering actually prevent model collapse long term?

Verifier-filtered synthetic retraining reduces variance and improves performance initially, but pulls model parameters toward the verifier's own knowledge center over time. Without perfect verifier reliability, early gains plateau and degrade as bias accumulates.

Where does mode collapse in language models really come from?

The research traces mode collapse to annotators' systematic preference for familiar text, a cognitive bias baked into training data. Verbalized Sampling, a training-free prompting method, restores 1.6–2.1× diversity by asking models to articulate probability distributions over responses.

Why do language models collapse into generic templates?

When within-prompt reward variance is low, task gradients weaken and regularization dominates, pushing policies toward generic outputs. SNR-Aware Filtering—selecting high-variance prompts before updates—recovers performance across tasks and scales.

Show all 6 sources
What makes synthetic data work across different domains and models?

Research shows no single optimal recipe for synthetic data generation. The impact of data properties like complexity and diversity varies by domain, model, use case, and scale, making explainable, flexible control more valuable than one-size-fits-all methods.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.