SYNTHESIS NOTE
Topics›Reasoning Critiques›this note

Does RL training collapse format diversity in pretrained models?

Exploring whether RL fine-tuning systematically selects one output format from pretraining while suppressing others, and how this selection mechanism drives performance gains.

Synthesis note · 2026-02-22 · sourced from Reasoning Critiques

A study with full pretraining transparency (models pretrained from scratch on known open datasets) reveals a striking structural pattern: RL fine-tuning does not simply improve reasoning — it systematically selects for and amplifies a single format from the pretraining mixture while collapsing all others.

The mechanism: early in RL training (within the first epoch), the model shifts toward generating outputs in the format of one specific distribution — code-like formats for smaller models, natural language formats for larger models. This transition coincides with the largest accuracy gain, suggesting the selection of a dominant format is what drives improvement, not a gradual enhancement across all formats.

Key findings:

This is distinct from Does policy entropy collapse limit reasoning performance in RL? in an important way. Entropy collapse describes diversity reduction within an output distribution. The echo chamber finding describes distribution selection: RL picks one distribution and amplifies it at the expense of all others. It is a format-level convergence, not just a diversity-level collapse.

The implication for practitioners: RL fine-tuning results depend on what the pretraining data mixture looks like, but this dependence is largely hidden when starting from existing pretrained models whose training data is proprietary. The performance gains attributed to RL algorithms may partially reflect which pretraining distribution was selected, not algorithmic superiority.

Inquiring lines that read this note 333

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do models learn from self-generated outputs without cascading failures? Can AI systems achieve real improvement without external human feedback? Does preference optimization undermine conversational grounding in language models? How does diversity prevent model convergence on superficial patterns? How do curriculum design and feedback approaches affect model learning? When do simpler collaborative filtering approaches outperform complex LLM recommenders? What prevents LLMs from applying their reasoning knowledge to improve outputs? What are the fundamental limits of prompting for language models? Why do LLM research ideation systems generate novelty but lack diversity? How do neural networks learn compositional structure from training? How do training data quality and composition affect downstream model performance? What explains the gap between benchmark scores and true reasoning capability? How does RLHF training shape models to prioritize agreement over accuracy? Which reinforcement learning modifications most improve dialogue quality in language models? Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones? Can base models hide emergent misalignment through alignment training? How does optimization for reward create emergent misalignment in language models? How does policy entropy collapse limit scaling of reasoning-focused reinforcement learning? What capabilities differentiate diffusion from autoregressive language models? Does AI deployment reduce or exacerbate workplace inequality and income instability? Why do training associations persist despite contradictory contextual information? Can minimal training unlock latent reasoning already present in base models? Does training data format shape model reasoning more than domain content? Can iterative DPO substitute for online RL in studying misalignment? Can readers reliably distinguish AI-written text from human writing? How does fine-tuning trade off accuracy against reasoning quality? Can smaller specialized models match frontier models on key metrics? How effectively can test-time voting aggregate diverse reasoning samples? How does model capacity affect learning performance on diverse downstream tasks? Can AI agents improve their skills through accumulated experience and reuse? Can AI systems evade safety evaluations through reasoning manipulation? What makes reasoning traces effective supervision even when they're incorrect? What prediction granularity best trains models to generate reliable reasoning? Do persona-based approaches introduce systematic biases in user simulation? Can latent reasoning match or exceed explicit reasoning performance? Can persona profiles improve LLM prediction accuracy and consistency? Does pretraining establish the ceiling for what reward learning can improve? How do reward signal properties affect model reasoning and safety? Do accumulated memories help or hurt continual learning in models? What limits recursive self-improvement in autonomous AI systems? How do reward models systematically fail to represent diverse human preferences? What structural patterns sustain successful multi-turn dialogue and prevent breakdown? Can code harness improvements rival direct model scaling for capability? When should retrieval systems decide to fetch new information? Does intelligent routing among smaller models outperform training larger models? Why do models reveal hidden associations despite concealment attempts? Can mechanistic interpretability methods reliably reveal what models actually know? How can persistent memory architectures preserve information across ultra-long contexts? How does decomposing tasks into separate stages affect reasoning quality and safety? What limits language model accuracy in evaluating ideas? Do single-axis benchmarks accurately measure agent capability for real deployment? How do real-world evaluations reveal AI capabilities that benchmarks hide? Can models develop genuine introspective capability, or only mimic it? Can recurrent computation unlock reasoning capabilities that fixed-depth models cannot? What gaps exist between benchmark performance and real deployment outcomes? Do evolved harnesses learn transferable strategies or task-specific optimization artifacts? Why do language models hallucinate and how can we prevent it? Can external verification systems adequately replace learned reasoning in AI outputs? How does scaling reasoning capabilities affect models' appropriate abstention behavior? Do language models reason through disagreement or only accommodate it? Does AI assistance help or harm professional skill development? Can inference-time computation adaptively substitute for static model capacity? Can AI research automation sustain progress through accelerating feedback loops? How reliably can humans and AI detectors identify machine-generated text? Do language models encode knowledge that influences generation, or primarily imitate surface patterns?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
16 direct connections · 154 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

rl post-training converges on a single dominant pretraining distribution format, suppressing all others