Does RL training collapse format diversity in pretrained models?
Exploring whether RL fine-tuning systematically selects one output format from pretraining while suppressing others, and how this selection mechanism drives performance gains.
A study with full pretraining transparency (models pretrained from scratch on known open datasets) reveals a striking structural pattern: RL fine-tuning does not simply improve reasoning — it systematically selects for and amplifies a single format from the pretraining mixture while collapsing all others.
The mechanism: early in RL training (within the first epoch), the model shifts toward generating outputs in the format of one specific distribution — code-like formats for smaller models, natural language formats for larger models. This transition coincides with the largest accuracy gain, suggesting the selection of a dominant format is what drives improvement, not a gradual enhancement across all formats.
Key findings:
- The dominant distribution is typically the most performant — RL selects for the format in which the base model is already strongest
- Scale-dependent bias — smaller models favor simpler, code-like formats; larger models shift toward natural language
- The amplification degree depends on KL penalty — looser KL constraints produce more extreme format collapse
- RL does not always favor the most common distribution — pretraining proportions predict which distribution "wins" only sometimes
This is distinct from Does policy entropy collapse limit reasoning performance in RL? in an important way. Entropy collapse describes diversity reduction within an output distribution. The echo chamber finding describes distribution selection: RL picks one distribution and amplifies it at the expense of all others. It is a format-level convergence, not just a diversity-level collapse.
The implication for practitioners: RL fine-tuning results depend on what the pretraining data mixture looks like, but this dependence is largely hidden when starting from existing pretrained models whose training data is proprietary. The performance gains attributed to RL algorithms may partially reflect which pretraining distribution was selected, not algorithmic superiority.
Inquiring lines that read this note 333
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do models learn from self-generated outputs without cascading failures?- What happens when models train on AI-generated content recursively?
- Why does online RL succeed where supervised training fails for self-correction?
- Why does self-generated training data outperform externally sourced data?
- Why does asymmetric self-play create naturally calibrated difficulty better than fixed curricula?
- What failure modes emerge when model-generated content trains on itself iteratively?
- Can the serving loop itself become the primary training data source?
- How does self-distillation differ from standard fine-tuning approaches?
- Can unsupervised confidence-based training scale to domains beyond human evaluation reach?
- What causes irreversible model collapse when training on model-generated content?
- Does self-generated training data reduce a model's capability diversity?
- Can self-training drift be prevented by applying student compatibility filtering?
- How does adversarial collapse threaten unsupervised self-play skill construction?
- Do correlated human errors prevent models from transcending their training sources?
- Does higher temperature sampling help bootstrapping loops generate more training data?
- Why do different AI models generate similar outputs independently?
- Which AI imaginaries dominate training data and shape system behavior most strongly?
- Can distillation help AIs scale their learned objectives across many copies?
- Why does RLHF alignment reduce the diversity of viewpoints in AI output?
- How does preference-based training compare to supervised fine-tuning for function calling?
- How does task-oriented fine-tuning compare to preference tuning methods?
- Why does preference tuning reduce diversity in code but increase it in creative tasks?
- What happens to model grounding when preference optimization increases effective diversity?
- When does RLHF reduce diversity and when does it preserve semantic variation?
- Why do preference-tuned models produce different diversity patterns in code versus creative writing?
- Does alignment compound cultural bias that started during pretraining?
- What role does rigid output format play in function calling failure modes?
- How do unstated constraints become invisible to training data distributions?
- When does natural context diversity reduce the need for explicit exploration?
- Can accelerated sampling techniques from image generation speed up evolutionary search?
- How does covariate diversity compare to the exploration assumptions of LinUCB?
- How do surface statistical regularities enable correct outputs while degrading robustness?
- Do different function-calling subtasks have different entropy profiles during training?
- What makes output convergence across models inevitable given input-side homogenization?
- Can structured output formats reduce instruction following degradation?
- Can diversity-aware RL objectives prevent format convergence?
- What role does KL penalty strength play in format selection?
- Can synthetic data generation balance all three QDC axes simultaneously?
- How do RL subnetworks identified from different random seeds compare?
- How does KL penalty strength affect the degree of format collapse during RL?
- Can RL format selection explain performance gains attributed to algorithmic improvements?
- Can shifting the accuracy metric itself eliminate the need for diversity post-processing?
- Why do queries with low cross-rollout variance produce degenerate gradients?
- Why do sparse parameter subsets enable full-rank learning in RL?
- Why should deep learning theory prioritize average-case over worst-case analysis?
- Can group-relative normalization be modified to resist shortcut trajectories?
- Why do six different RLVR algorithms converge on similar performance levels?
- How does probability mass concentration affect sampling diversity across model scales?
- Can experimental outcomes be reliably distilled into reusable insights?
- How do normalization and input injection control emergence of fixed points?
- How does mutual information between inputs and outputs differ from measuring raw diversity?
- Why is the fast non-parametric loop vulnerable to overfitting differently than model weights?
- How does tournament selection without gradients compare to gradient-based hyperparameter tuning?
- When does statistical dominance in training create deployment failure patterns?
- What training signals would models need to learn reciprocal common-ground construction?
- How do training objectives shape what a world model actually learns?
- Does the model learn depth-wise drift as an explicit strategy?
- How do different training objectives shift whether models over-predict or under-predict?
- How does distributional distance from pre-training relate to model difficulty?
- Why does curriculum learning with tight budgets beat fixed-budget approaches?
- How does behavioral fine-tuning differ from factual knowledge encoding in models?
- How should researchers evaluate whether correct model outputs reflect real structural learning?
- Can backward transfer measurements reliably predict optimal multi-task training order?
- Why does training order matter across different domain types?
- Does critique training improve exploration diversity during model training or only test time?
- Does weight decay directly cause contractive behavior near training examples?
- What happens to base model capabilities when you apply finetuning?
- Why does the order of training examples matter for what models learn?
- How much does pretraining quality affect the modularity of fine-tuned models?
- What features does a sample reinforce when it moves bands?
- What mechanisms cause overly hard samples to degrade prior model performance?
- Can trained models encode programs more complex than their data-generating process?
- What training regimes confound surface mechanisms with their actual causes?
- Can models generate their own training curriculum during offline dreaming?
- How do finetuning and pretraining improvements differ in their effects on model capabilities?
- Why do harness validators shape what models learn to emit?
- Does curriculum-based training keep small models perpetually at their learning edge?
- What trade-offs emerge between training objectives and model reliability?
- What distinguishes surface mechanisms from the training regimes that produce them?
- How do learning dynamics on one example shift predictions on other responses?
- How does the proxy pattern explain failures in RL-based safety training?
- Why does narrow training data produce broad harmful behavior patterns?
- Why does recontextualizing a behavior during training change whether models learn it?
- Why does fine-tuning function as character training rather than capability training?
- Can few-shot examples narrow generative diversity in creative tasks?
- Can this whole-artifact principle apply to other generative tasks?
- Can world models form from aggregated partial information across training distributions?
- What happens when a single loss function conflates representation learning with decision-making?
- What neural or architectural mechanism allows selective override of frequency effects?
- How do gradients flowing through both branches simultaneously reshape each component's role?
- What makes data augmentation an implicit form of contraction learning?
- How do encode-decode contractive biases create stable attractors in latent space?
- What solvable idealized settings reveal fundamental phenomena in realistic deep learning?
- How do weight visualizations reveal temporal structure in cyclic training?
- Can training order and structure shape what networks retain and learn?
- Does latent density emerge during pretraining from training data familiarity?
- Why do proprietary models improve with training while open-source models decline?
- What conditions make training diversity better than individual expert quality?
- Why does mixed instruction data sometimes hurt specific model capabilities?
- How does training data distribution determine what models can learn?
- How does training frequency distribution shape what models reliably retrieve?
- When should full-parameter post-training be used instead of LoRA adaptation?
- What creates the irreducible trade-off between quality and diversity in training data?
- How does diversity loss in synthetic data mirror tail distribution disappearance?
- How do quality, diversity, and complexity create different effects on downstream model performance?
- Why do certain tokens at certain difficulties drive most of RLVR's learning signal?
- At what point does output quality outweigh diversity value in synthetic data tasks?
- How much does diversity training cost in single-shot pass@1 performance?
- Why do unified models still inherit data-distribution biases from training?
- Does verbalized sampling preserve factual accuracy and safety during diversity gains?
- Can decoding-time prompting strategies fully replace diversity-focused training methods?
- Can data pruning and equal contribution be reconciled in optimal learning?
- How much performance is lost when converting pretrained checkpoints versus training from scratch?
- How do complexity and diversity affect model performance differently?
- Why does the same training data produce different gains across models?
- How do cyclic learning rates anti-correlate with weight decay to create diversity?
- Does unpredictable generalization from SDF become predictable at different training document scales?
- How should training distribution distance be defined when the policy evolves?
- Why does diversity in training data enable denoising rather than reinforce shared biases?
- Can synthetic data diversity preserve the transcendence effect or does it collapse?
- Do identical task structures mean repeated instances or new synthetic samples with same design?
- Can scaling up contradictory training data overcome unpredictable override effects?
- Can training data organization by capability outperform source or task-based mixing?
- Does teacher-style refinement of training data transfer equally to all student model distributions?
- Can complexity, diversity, and fidelity scale together in synthetic environments?
- How much does training data composition shape security model performance?
- Why does synthetic-only data degrade performance even when matched in size?
- How do quality and diversity in synthetic data interact with accumulation schedules?
- Can temperature-adaptive sampling reduce the sharpening tax during training?
- Why does diversity of training cases matter more than raw dataset size?
- How much RLVR improvement comes from benchmark data memorization?
- Why do internal representations differ when external performance matches?
- Can fine-tuning or RLHF alone solve the persona distortion problem?
- Does RLHF training suppress exploratory and qualifying language?
- Can RLHF training push models away from human-like lexical patterns?
- Why does better RLHF training fail to decouple polish from persona distortion?
- What's the difference between RLHF, RLVR, and RLCF as training paradigms?
- Why does post-training alignment create skew in simulated survey responses?
- Can RLHF training signals work as well for prose as they do for math and code?
- Can distillation methods extract directional guidance that scalar RL cannot access?
- Can messy multi-agent transcripts become better training data than clean outputs?
- Can proper scoring rules fix RLVR's degradation on disagreement prediction?
- Does environment stochasticity force models to generalize better across trajectory variations?
- What scaling properties emerge from RL training dynamics beyond verification?
- How does curriculum learning prevent instability in social-emotional RL training?
- How should multi-objective post-training balance competing behavioral goals?
- Does semantic diversity in output space compete with reward-component diversity?
- Can feedback loop frequency harm performance on finite task sets?
- Why does alternating RL training stabilize learning better than simultaneous updates?
- How does non-reasoning SFT prevent overfitting before RL training begins?
- Can in-context learning replicate the timing effects that RL teaches models?
- How does reinforcement learning compare to differentiable joint training for RAG?
- Can smaller models achieve domain expertise through focused RL training?
- How does RL compress reasoning path diversity during training?
- Does sparsity in RL arise from training on policy-distribution data?
- Why does RL improve sampling efficiency but not expand capability boundaries?
- What limits RLVR effectiveness beyond mathematical and coding domains?
- Does RLVR expand model capability or reorganize existing capability?
- Does format-based pretraining determine how models respond to reinforcement learning?
- Can explicitly optimizing for semantic diversity during RL training improve both quality and variation?
- Why do overtrained domains show different RL training outcomes than novel tasks?
- What makes supervised fine-tuning worsen RL exploration later?
- How does prolonged RL training differ from standard RLVR approaches?
- Why does supervised fine-tuning on diverse demonstrations expand exploration diversity compared to RL?
- How do verifier-free RL patterns differ from traditional RLHF approaches?
- Does RLVR teach new reasoning or activate existing pretraining capabilities?
- How do sparse parameter updates enable when-not-how training to work?
- Can single-problem fine-tuning match full RL pipeline reasoning gains?
- Why does reinforcement learning training degrade model calibration?
- Why does outcome-based RL specifically lose diversity during training?
- Can RL directly optimize attention distributions instead of text generation?
- How does imitation pretraining followed by RL exploration compare to either method alone?
- Why do reasoning gains from RL require models trained with headroom and edge-of-competence data?
- Why does exploration diversity behave differently under reinforcement learning versus supervised fine-tuning?
- How does pretraining quality versus quantity affect downstream RL gains?
- Can pretraining alone achieve chess performance without reinforcement learning?
- Can RL training on small verifiable tasks transfer to real-world AI research?
- Does alignment training create bidirectional instruction and response mappings?
- Does removing cognitive bias from training signals accidentally break what makes alignment work?
- Why does post-training suppress alignment faking in some models but amplify it in others?
- Can alignment training create systematic blind spots in threat detection systems?
- Does pretraining poisoning at scale persist through instruction alignment?
- What alignment procedures cause different models to share the same output distribution?
- How does upstream value embedding differ from downstream alignment patches?
- Can mechanistic interpretability tools decode the biases alignment training conceals?
- How do alignment priors drive similar outputs across different models?
- Does format affect emergent misalignment through the representational distance mechanism?
- Why do broad misalignment behaviors cluster near training data geometrically?
- Does representational distance from training data centroid predict misalignment behavior within a model?
- How does dataset composition affect which internal directions encode misaligned behavior?
- How does post-training affect alignment faking across different model architectures?
- Can training or alignment changes explain the regression in frontier models?
- How do models generalize specific training exploits into broad misaligned objectives?
- Can models trained with RL on pretraining data avoid reward hacking seen in RLHF?
- Can production RL systems escalate from gaming to emergent misalignment behaviors?
- Why does RLVR increase token entropy while decreasing answer diversity?
- What distinguishes training-time entropy collapse from test-time variance inflation?
- How does inference variance differ from training entropy collapse?
- Is distribution selection during RL the same compression mechanism as entropy collapse?
- What role do high-entropy minority tokens play in RLVR?
- How does representational convergence differ from policy entropy collapse in iterative training?
- How does entropy loss enable exploration beyond a single training example?
- How does on-policy entropy recognition differ from training-time entropy collapse?
- Can entropy regularization or critique models prevent search strategy collapse during RL training?
- Does post-training collapse policy entropy more than base model sampling?
- Why does prolonged RL with entropy control beat base models at all pass@k levels?
- Can feature disentanglement in gesture synthesis generalize to completely unseen voice distributions?
- What specific optimizations from LLM training transfer back to encoder models?
- How do early layers preserve unbiased information while late layers conform?
- Why does context information fail to override prior training associations?
- How much can mitigation techniques like augmentation reduce priming without harming learning?
- Does foundational model training or user priors more strongly shape final outputs?
- Why does consistency training make models resistant to prompt perturbations?
- Do negative constraints require fundamentally different training signals than positive instructions?
- How much does pretraining contribute to ToM performance versus task-specific training?
- Why does combining reasoning distillation with RLVR outperform either training stage alone?
- Why does training data format matter more than domain content?
- Why does training data format matter more than its domain content?
- Does training data format shape model reasoning more than domain content?
- How does training data format shape whether models reason in parallel or sequentially?
- Does training data format matter more than who generates it?
- How does training data format shape which reasoning patterns emerge in models?
- Does training data format determine whether models collapse entropy or inflate variance?
- Can training format itself shape what reasoning strategy a model learns?
- Why do zero-advantage rollouts destabilize training beyond just wasting compute?
- Why does gradient discarding limit standard policy clipping?
- Can trust region constraints prevent the sample inefficiency problems of RLHF?
- What misaligned behaviors does iterative DPO produce compared to online RL?
- What theoretical argument connects iterative DPO dynamics to online RL learning?
- Why does NLI fine-tuning amplify frequency bias instead of teaching inference?
- Why does fine-tuning change how models process retrieved context?
- Does fine-tuning on NLI tasks reduce or amplify frequency bias?
- Why does fine-tuning improve some capabilities while degrading others?
- Why does fine-tuning fail to remove temporal contamination from pretraining?
- Do instruction-tuned models learn tasks or just output format distributions?
- Why does eliminating proxy-model filtering improve reasoning emergence in pretraining?
- How should skill libraries coordinate with gradient-based weight optimization?
- Where does skill extraction fail compared to genuine model adaptation?
- Do text-space skills transfer learning across different frontier models?
- Does finetuning facts into weights overwrite existing model capabilities?
- How do ensemble methods apply within a single model?
- Why do production systems optimize for three model classes instead of foundation models?
- Why do production teams choose expensive frontier models over fine-tuning?
- Can specialized components replace single fully-trained models in deployment?
- Why does capability saturation and diversity saturation occur at different scales?
- Why does low temperature sampling extract consensus from diverse training data?
- How does training-time voting differ from inference-time majority voting over samples?
- Why do self-consistency methods fail where pretraining bias is strongest?
- What signals detect when consensus training is silently degrading performance?
- What capabilities actually require massive scale versus specialized training regimes?
- Why do smaller and larger models converge on different output formats?
- Does fine-tuning actually change model capabilities or only output distribution?
- Can models converge on similar experience descriptions across different architectures?
- What output distribution properties make smaller models better for wide sampling?
- Why should scaling laws be understood as properties of data distribution rather than training in general?
- Which finetuning method works best across different task and data regimes?
- How do task frequency and complexity interact with model capacity during training?
- Can dense models match specialized architectures by mixing data better?
- Why does model×environment interaction dominate recognition variance?
- How does model scale affect the crossover point between base and post-trained performance?
- How does mutual shaping through diverse training compare to population-level diversity effects?
- Can skill libraries prevent redundant narrow artifacts from proliferating?
- Does a tight learning-rate bound prevent skills from escaping poor starting points?
- Can RL-trained policies outperform text-space optimizers for evolving skill repositories?
- Does reasoning trace style explain why RL post-training improves model reasoning?
- Why does long CoT training optimize for structural coherence over content correctness?
- Does training on model-generated correction traces actually work?
- Why did prior multi-token prediction methods fail during fine-tuning?
- Do high-entropy RLVR tokens correspond to MI-peak tokens during inference?
- How could persona vector tracking complement multi-turn RL for earlier drift detection?
- Does the Assistant Axis exist in pre-trained models before instruction tuning?
- Does pre-training encode personality patterns that fine-tuning later activates?
- How do pretraining priors shape what models invent about users?
- Which recipe choices determine the asymptotic ceiling in RL training?
- How does behavior cloning reduce complexity before RL training in rerankers?
- How does pretrained knowledge constrain what adaptation strategies can achieve?
- What makes pretraining composition more important than reward engineering?
- How do RL training and base models differ in creating MI peaks?
- How do self-evolving curricula help RL break beyond base model capability boundaries?
- How does post-training shift models from passive prediction to on-policy action?
- Does RL training activate latent meta-learning capacity or create it from scratch?
- Does the pretrained model prior limit RL search capability more than the optimization algorithm itself?
- What pretraining formats encode latent reasoning strategies that RLVR can surface?
- Why does the pretrained prior determine the exploration ceiling?
- How does pretraining determine what RL can later teach a model?
- How does model scale affect anticipatory behavior in structured training?
- How does active selection of training content differ from random reinforcement sampling?
- What role does pretraining play in distinguishing system capability from deployed behavior?
- Will future training data teach models to connect awareness with gaming?
- Does negative reinforcement alone achieve what full RL training accomplishes?
- How do reward signals in RLVR interact with pretraining biases?
- Does careful reward engineering matter if pretraining determines RLVR effectiveness?
- Why do pretrained model priors reduce the usefulness of retrieved experience?
- How do trained weights differ from a stored library or text?
- Can the joint-training principle extend beyond memorization and generalization pairs?
- How does KL regularization prevent both forgetting and adaptation loss?
- How does in-weights adaptation create spurious forgetting in models?
- Do sample-level similarities between pretraining and downstream tasks explain the frequency effect?
- What happens when you project the same model onto different harnesses?
- How does editing the harness layer differ from updating model weights?
- How do pre-training and distillation enable minimal routing signals to work?
- Can routing signals organize training data into a meaningful curriculum automatically?
Related concepts in this collection 6
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does policy entropy collapse limit reasoning performance in RL?
As reinforcement learning models become more confident in their policy choices, entropy drops and performance plateaus. Can we identify and counteract this bottleneck to sustain scaling?
entropy collapse is the within-distribution consequence; this note is the between-distribution mechanism
-
Why do reasoning models fail differently at training versus inference?
Reasoning models exhibit two distinct failure modes—entropy collapse during training and variance inflation during inference—that appear unrelated but may share underlying causes. Understanding these dual problems could reveal whether separate or unified solutions are needed.
adds a third layer: not just entropy collapse and variance inflation, but distribution selection
-
Can simple rewards alone teach complex domain reasoning?
Does reinforcement learning on difficult problems with basic accuracy rewards produce sophisticated reasoning strategies without explicit chain-of-thought training? This challenges assumptions about what domain AI models need to learn effectively.
emergence through RL looks different when the pretraining mixture is known: it's partly selection, not purely emergence
-
Does RL improve domain reasoning by adding knowledge or removing it?
When reinforcement learning improves reasoning in specialized domains like medicine, is it teaching models new facts or preventing them from using wrong ones? Understanding this distinction matters for how we design RL training.
pruning operates within the selected distribution; this note shows which distribution gets to keep its knowledge
-
Does reinforcement learning squeeze exploration diversity in search agents?
Investigates whether RL training narrows the behavioral diversity of search agents the same way it does in reasoning tasks. Understanding this mechanism could reveal whether entropy collapse is fundamental to RL or domain-specific.
confirms the echo chamber dynamic is domain-general: RL squeezes search strategy diversity just as it selects a single pretraining format — format selection and within-format entropy collapse are two levels of the same RL compression
-
Why does RLVR training narrow a model's problem solving ability?
RLVR's on-policy constraint may force models to exploit known reasoning paths rather than explore new ones, potentially shrinking their effective problem-solving scope. Understanding this mechanism could reveal how to design better exploration incentives in language model reasoning.
capability boundary collapse is the downstream consequence of format selection: when RL selects one dominant distribution, problems solvable only through suppressed formats become unreachable
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
- The Art of Scaling Reinforcement Learning Compute for LLMs
- Sharpening Tax in Post-Training
- Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs
- Understanding Reasoning from Pretraining to Post-Training
- Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
- 1000 Layer Networks for Self-Supervised RL: Scaling Depth Can Enable New Goal-Reaching Capabilities
- Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Original note title
rl post-training converges on a single dominant pretraining distribution format, suppressing all others