INQUIRING LINE

Why does training an AI with the right 'exploration brakes' make it better at literally every try — not just its first guess?

Why does prolonged RL with entropy control beat base models at all pass@k levels?

This explores why one line of RL research finds trained models beating their base models at every pass@k level (the chance that at least one of k attempts is correct), when RL usually gives a model better first tries but leaves it no better, or worse, when it gets many tries.


This explores why RL training sometimes makes a model better at every sampling budget, not just on its first attempt. To see why that's surprising, start with the usual pattern: RL tends to sharpen a model rather than teach it anything new. Training pushes the model's probability toward answers it could already reach, so pass@1 goes up. But the range of approaches it will try shrinks, and when you let it make many attempts (high k), the base model often catches up or pulls ahead. The paper that breaks this pattern is Can reinforcement learning discover reasoning strategies base models cannot?. It credits three ingredients: KL control (a penalty that keeps the model from drifting too far from where it started), periodic policy resetting, and training on varied non-math tasks where the base model has no well-worn habits to fall back on.

Why entropy control matters becomes clear once you see what goes wrong without it. Does policy entropy collapse limit reasoning performance in RL? finds a simple law: performance rises as policy entropy (roughly, how varied the model's choices are) falls, and it levels off once entropy nears zero. A model that has stopped exploring can't find anything new, so its ceiling is fixed by what it had already found. The same narrowing shows up in several forms across the corpus. RL locks onto one output format from pretraining and suppresses the rest within the first epoch (Does RL training collapse format diversity in pretrained models?). It squeezes the variety of strategies search agents use (Does reinforcement learning squeeze exploration diversity in search agents?). When rewards barely vary between attempts on the same prompt, it can even fall back to generic, input-agnostic templates (Why do language models collapse into generic templates?). Low pass@k for base models at high k is what all of this narrowing looks like from the outside.

This suggests that 'prolonged' only pays off when exploration survives. Long training without entropy control just collapses more thoroughly. With entropy control, something else can happen. Does RL training follow a predictable two-phase learning sequence? finds that RL first spends its effort on getting execution right. Only later does the bottleneck move to strategic planning, and in that phase the entropy of planning tokens rises. If entropy has already collapsed by the time phase two arrives, the model never gets to explore strategies. Keeping it alive is what lets the model reach the phase where new approaches show up.

KL control may also help for a less obvious reason: it keeps the model trainable. Does staying close to the base model preserve learning ability? shows that models staying closer to their base distribution keep their ability to learn new tasks, while models that drift further stall when the domain changes. Policy resetting fits the same logic, since it periodically returns the model to a state where it can still learn. And the change RL makes may be more targeted than it looks: Does reinforcement learning update only a small fraction of parameters? finds that RL touches only a small, consistent subset of parameters. Constraining drift doesn't have to mean constraining what the model can learn.

One caveat. The corpus has one direct demonstration of the all-pass@k result, and that paper points out it is strongest on domains where the base model 'lacks established patterns.' On math, where base models already have strong habits, sharpening may still be most of what RL does. So the question isn't whether RL can teach new reasoning in general. It's whether you keep exploration alive long enough, on problems unfamiliar enough, for anything new to be found.


Sources 8 notes

Can reinforcement learning discover reasoning strategies base models cannot?

RL-trained models outperform base models across all pass@k levels when trained with KL control, policy resetting, and non-mathematical tasks. This shows RL can expand capability boundaries, not just optimize sampling efficiency, especially on domains where base models lack established patterns.

Does policy entropy collapse limit reasoning performance in RL?

Empirical law R = -a·exp(H) + b shows performance saturates when policy entropy approaches zero. Interventions like Clip-Cov, KL-Cov, and GPPO preserve exploratory capacity by managing entropy reduction during training.

Does RL training collapse format diversity in pretrained models?

Controlled experiments show RL consistently amplifies one format distribution from pretraining within the first epoch while collapsing alternatives. The winning format depends on model scale, not necessarily performance, and is largely hidden when starting from proprietary pretrained models.

Does reinforcement learning squeeze exploration diversity in search agents?

RL training compresses behavioral diversity in search agents through the same entropy collapse mechanism documented in reasoning—policies converge on narrow reward-maximizing strategies. SFT on diverse demonstrations preserves exploration breadth, suggesting diversity-preservation techniques are essential for RL search scaling.

Why do language models collapse into generic templates?

When within-prompt reward variance is low, task gradients weaken and regularization dominates, pushing policies toward generic outputs. SNR-Aware Filtering—selecting high-variance prompts before updates—recovers performance across tasks and scales.

Show all 8 sources
Does RL training follow a predictable two-phase learning sequence?

Across eight models, RL training consistently shows a first phase where execution correctness drives learning, followed by a second phase where strategic planning becomes the bottleneck. Planning token entropy increases while execution entropy stabilizes, and concentration of optimization on planning tokens yields significant performance gains.

Does staying close to the base model preserve learning ability?

FST-trained models stay up to 70% closer to their base distribution than parameter-only RL, and this reduced drift preserves the model's ability to learn subsequent tasks effectively. Parameter-only approaches stall when task domains change, while low KL drift enables sustained adaptation.

Does reinforcement learning update only a small fraction of parameters?

Across seven RL algorithms and ten LLM families, RL induces intrinsic parameter sparsity of 5–30% without explicit regularization. Critically, these sparse updates are nearly full-rank and nearly identical across random seeds, indicating structural rather than arbitrary parameter selection.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.