INQUIRING LINE

Does synthetic training data need to be good, varied, or just mixed in alongside real data the right way?

How do quality and diversity in synthetic data interact with accumulation schedules?

This explores whether the way synthetic data is fed back into training (replacing real data versus piling up alongside it) changes how its quality and diversity matter, and whether the two problems are really one.


This explores whether the way synthetic data enters training (swapped in for real data, or added alongside it) changes how much its quality and diversity matter. The collection has no single paper that tests all three together. Read side by side, though, the notes suggest something unexpected: the main reason the schedule matters may be diversity. The clearest result on schedules is Does model collapse depend on how we schedule training data?. If each generation's synthetic output replaces the real data, test error grows without limit. If synthetic data accumulates on top of the original real corpus, error stays bounded. This held for language models, diffusion models and VAEs. One plausible reading is that the real data works as a diversity anchor: it keeps the full spread of the original distribution present, so nothing can quietly drop out.

That fits a separate line of work that splits "good synthetic data" into parts. How do quality, diversity, and complexity affect synthetic data differently? finds that quality mainly helps a model on data like what it has seen. Diversity is what lets it cope with situations it hasn't seen. Complexity helps both. Because common evaluations fold all three into one quality score, self-improvement loops tend to keep raising quality while losing diversity, and that loss is hard to reverse. Put the two notes together and you get a hypothesis the corpus suggests but doesn't directly test. Replacement schedules are dangerous because each round filters for quality and lets the variety shrink. Accumulation schedules survive because the original variety is never thrown away.

The same pattern of variety vanishing without anyone noticing shows up in other places. Does synthetic content in search results hide ecosystem decay? shows a search corpus that becomes 67% synthetic. Over 80% of retrieved results then come from synthetic sources, yet answer accuracy stays high. The monoculture shows up only when it breaks. Inside training, Does RL training collapse format diversity in pretrained models? finds that RL amplifies one output format from pretraining within the first epoch and suppresses the rest. The pattern repeats: optimizing for 'good' answers shrinks variety, and the usual metrics don't flag it. One response is to reward diversity directly. Can diverse mediocre traces outperform redundant expert traces? scores a whole set of outputs together, and finds that varied mediocre traces can beat a set of near-identical strong ones.

If diversity is what's at risk, the next question is how to generate it on purpose instead of hoping real data keeps supplying it. Can we generate synthetic data without any seed examples? (Simula) handles broad coverage with a topic taxonomy and handles complexity separately, so quality, diversity and complexity can each be tuned. Can synthetic dialogues become realistic through layered diversity? builds realistic dialogue variety from several layers: subtopic, personality and context. What makes synthetic data work across different domains and models? adds a caveat: the right balance changes with domain, model and scale, so no single schedule or mix will be best everywhere.

One more idea changes the question. Should we treat LLM outputs as real empirical data? argues that model-generated text reflects the model's learned patterns and the prompt it was given, not new observations about the world. It should count only through an explicit trust weight. Seen this way, "accumulate, don't replace" is a rough version of a better rule: never let synthetic data outweigh the real evidence it came from. The collection does not yet hold a study that varies the accumulation schedule and synthetic-data diversity at the same time. That gap is worth flagging.


Sources 9 notes

Does model collapse depend on how we schedule training data?

Replacing real data with synthetic data causes unbounded test error growth, but accumulating synthetic data alongside the original real corpus keeps error bounded across model architectures and sizes. The mechanism is proven analytically in a linear-regression framework and confirmed empirically on language models, diffusion models, and VAEs.

How do quality, diversity, and complexity affect synthetic data differently?

Quality drives in-distribution generalization, diversity enables out-of-distribution generalization, and complexity strengthens both. Current evaluation methods collapse these into a single quality metric, causing self-improvement loops to degrade through irreversible diversity loss.

Does synthetic content in search results hide ecosystem decay?

When 67% of a corpus becomes synthetic, over 80% of retrieved results shift to synthetic sources while answer accuracy remains high, masking the loss of source diversity. This creates fragility: high accuracy resting on a monoculture collapses when that monoculture is poisoned.

Does RL training collapse format diversity in pretrained models?

Controlled experiments show RL consistently amplifies one format distribution from pretraining within the first epoch while collapsing alternatives. The winning format depends on model scale, not necessarily performance, and is largely hidden when starting from proprietary pretrained models.

Can diverse mediocre traces outperform redundant expert traces?

SPIRAL shifts RL reward from individual traces to sampled sets, optimizing for complementarity rather than per-trace accuracy. Diverse mediocre traces outperform redundant strong ones because aggregators need raw material to arbitrate, not confirmation.

Show all 9 sources
Can we generate synthetic data without any seed examples?

Simula separates global coverage from local diversity, using taxonomy construction for coverage and agentic refinement for complexity. This architecture makes all three desiderata—quality, diversity, complexity—controllable simultaneously without requiring seed data.

Can synthetic dialogues become realistic through layered diversity?

Research shows that realistic synthetic dialogues require three multiplicative layers: subtopic specificity, Big Five persona variation, and 11 contextual characteristics via Chain of Thought reasoning. This structured approach captures 90.48% of in-domain dialogue performance.

What makes synthetic data work across different domains and models?

Research shows no single optimal recipe for synthetic data generation. The impact of data properties like complexity and diversity varies by domain, model, use case, and scale, making explainable, flexible control more valuable than one-size-fits-all methods.

Should we treat LLM outputs as real empirical data?

Foundation Priors framework shows that LLM-generated text reflects the model's learned patterns and user's prompt choices, not ground truth. Such outputs should only influence inference through explicitly parameterized trust weights, not be treated as equivalent to real evidence.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.