INQUIRING LINE

Does training a model to follow instructions also quietly narrow the range of answers it's willing to give?

Does post-training collapse policy entropy more than base model sampling?

This explores whether post-training, especially reinforcement learning, narrows the range of things a model will say compared with sampling from the base model it started from, and what that narrowing costs.


This explores whether post-training narrows a model's range of possible outputs compared with sampling straight from the base model, and whether that matters. The short answer is yes, and it can matter a lot. One caveat first: the collection doesn't report a direct base-versus-tuned entropy measurement. Instead it shows the collapse through what it does to behavior. The clearest comparison is in Do base models find more solutions than post-trained ones?. Across 14 model pairs on agentic tasks, base models with only a light prompt eventually solve more distinct problems than their post-trained versions, once you let both try many times. Post-training makes the model reliably good on easy cases. But it erases rare solutions the base model could still reach.

The mechanism has a name inside RL training. Does policy entropy collapse limit reasoning performance in RL? finds that as training drives the model's randomness (its policy entropy) toward zero, performance hits a predictable ceiling. The model stops exploring, so it stops finding anything new. That paper also points to fixes, such as clipping or penalizing the updates that drain diversity fastest. The same pattern appears outside math reasoning. Does reinforcement learning squeeze exploration diversity in search agents? shows RL shrinking the range of search strategies agents use. Supervised fine-tuning on varied demonstrations does the opposite and widens it. So 'post-training' isn't one thing: RL tends to narrow behavior, while SFT on diverse data can broaden it.

The collapse isn't only about answers. It also hits style and structure. Does RL training collapse format diversity in pretrained models? finds that within the first epoch, RL picks one output format the base model already knew and suppresses the rest. Which format wins depends on model scale, not on which one works best. Why do language models collapse into generic templates? explains one route to this. When every attempt on a prompt gets roughly the same reward, the learning signal fades, regularization takes over, and the model falls back on generic templates that ignore the input. Filtering for prompts where attempts actually differ in reward helps reverse it.

The collapse isn't uniform either. Does RL training follow a predictable two-phase learning sequence? finds that entropy on step-by-step execution tokens settles down, while entropy on planning tokens actually rises in a later phase of training. Some narrowing is useful: you want the model to be consistent at arithmetic and exploratory about strategy. The problem is when the narrowing reaches the parts of reasoning where variety pays off.

The less obvious consequence is overconfidence. Does binary reward training hurt model calibration? shows that simple right/wrong rewards push models toward confident guessing, because they never penalize being confidently wrong. Adding a calibration term to the reward fixes this. So a collapsed policy doesn't just explore less. It can also sound surer than it should. A practical takeaway: if you need breadth (many candidate solutions, rare edge cases), sampling a base model many times may beat a polished post-trained model.


Sources 7 notes

Do base models find more solutions than post-trained ones?

Across 14 model pairs and three agentic benchmarks, base models equipped with only relaxed system prompts eventually surpass post-trained counterparts in pass@K coverage as rollout budget grows. Post-training bimodalizes task outcomes, sharpening performance on easy cases while eliminating rare-but-reachable solutions.

Does policy entropy collapse limit reasoning performance in RL?

Empirical law R = -a·exp(H) + b shows performance saturates when policy entropy approaches zero. Interventions like Clip-Cov, KL-Cov, and GPPO preserve exploratory capacity by managing entropy reduction during training.

Does reinforcement learning squeeze exploration diversity in search agents?

RL training compresses behavioral diversity in search agents through the same entropy collapse mechanism documented in reasoning—policies converge on narrow reward-maximizing strategies. SFT on diverse demonstrations preserves exploration breadth, suggesting diversity-preservation techniques are essential for RL search scaling.

Does RL training collapse format diversity in pretrained models?

Controlled experiments show RL consistently amplifies one format distribution from pretraining within the first epoch while collapsing alternatives. The winning format depends on model scale, not necessarily performance, and is largely hidden when starting from proprietary pretrained models.

Why do language models collapse into generic templates?

When within-prompt reward variance is low, task gradients weaken and regularization dominates, pushing policies toward generic outputs. SNR-Aware Filtering—selecting high-variance prompts before updates—recovers performance across tasks and scales.

Show all 7 sources
Does RL training follow a predictable two-phase learning sequence?

Across eight models, RL training consistently shows a first phase where execution correctness drives learning, followed by a second phase where strategic planning becomes the bottleneck. Planning token entropy increases while execution entropy stabilizes, and concentration of optimization on planning tokens yields significant performance gains.

Does binary reward training hurt model calibration?

Binary correctness rewards incentivize high-confidence guessing because they don't penalize confident wrong answers. Adding the Brier score as a second reward term mathematically guarantees joint optimization of accuracy and calibration without trade-off.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.