SYNTHESIS NOTE
Topics›Deep Research›this note

Does reinforcement learning squeeze exploration diversity in search agents?

Investigates whether RL training narrows the behavioral diversity of search agents the same way it does in reasoning tasks. Understanding this mechanism could reveal whether entropy collapse is fundamental to RL or domain-specific.

Synthesis note · 2026-02-21 · sourced from Deep Research

The "RL Squeezes, SFT Expands" paper studies search agents trained with RL versus SFT and finds the same pattern that the reasoning literature documented: RL training compresses the diversity of behaviors the agent explores (squeezes), while SFT on diverse demonstrations expands it. Since Does policy entropy collapse limit reasoning performance in RL?, and since this paper shows the same dynamic in search RL, entropy collapse is not a quirk of reasoning training — it is a property of RL training at large.

The mechanism is the same in both domains: RL rewards the policy for high-reward outputs and penalizes low-reward ones. Over training, the policy concentrates probability mass on the reward-maximizing region of its action space. In reasoning, this means converging on a narrow set of reasoning patterns. In search, it means converging on a narrow set of query strategies. Both reduce the agent's ability to explore novel approaches to hard problems.

SFT has the opposite effect because it trains on human demonstrations or diverse synthetic completions — the diversity of the training set is preserved in the policy. The tradeoff is that SFT cannot generalize beyond its demonstrations in the same way RL can.

This finding has practical implications for DR agent design: RL-trained search agents need explicit diversity mechanisms (entropy regularization, diverse reward models, periodic SFT refreshes) or they will converge on query templates that work well on average but fail on distribution shift. The same Do critique models improve diversity during training itself? remedy applies — external critique prevents the RL agent from collapsing to a narrow search strategy.

Inquiring lines that read this note 168

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Does AI assistance help or harm professional skill development? Why do LLM research ideation systems generate novelty but lack diversity? When do multi-agent systems improve over single frontier models? How do reward signal properties affect model reasoning and safety? How does policy entropy collapse limit scaling of reasoning-focused reinforcement learning? How does fine-tuning trade off accuracy against reasoning quality? Can AI agents improve their skills through accumulated experience and reuse? How does decomposing tasks into separate stages affect reasoning quality and safety? How does diversity prevent model convergence on superficial patterns? What limits recursive self-improvement in autonomous AI systems? How do neural networks learn compositional structure from training? Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones? What limits language model accuracy in evaluating ideas? Can smaller specialized models match frontier models on key metrics? Can language models reliably simulate personas and predict behavior? What are the fundamental limits of prompting for language models? How do training data quality and composition affect downstream model performance? Can latent reasoning match or exceed explicit reasoning performance? How does model capacity affect learning performance on diverse downstream tasks? How effectively can test-time voting aggregate diverse reasoning samples? Do persona-based approaches introduce systematic biases in user simulation? Does pretraining establish the ceiling for what reward learning can improve? How do agents learn to distinguish valuable feedback from noise? Can base models hide emergent misalignment through alignment training? Can AI systems discover fundamental improvements to their own architectures? How do curriculum design and feedback approaches affect model learning? When does parallel reasoning outperform sequential reasoning with the same token budget? Why do autonomous agents misreport success on failed actions? Should agents compress episodic memory or retain raw interaction histories? Does preference optimization undermine conversational grounding in language models? Which reinforcement learning modifications most improve dialogue quality in language models? Can inference-time computation adaptively substitute for static model capacity? Does scaling reasoning capability create fundamental tradeoffs in control and reliability? Do evolved harnesses learn transferable strategies or task-specific optimization artifacts? Can iterative DPO substitute for online RL in studying misalignment? How does awareness of evaluation context influence model behavior? How do AI systems determine and balance multiple competing objectives? Does AI-assisted research sacrifice exploration breadth for productivity gains? Do single-axis benchmarks accurately measure agent capability for real deployment? Are AI-generated articles systematically disadvantaged in search ranking and user engagement?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 116 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

rl training for search agents squeezes exploration diversity while sft expands it — the same entropy collapse dynamic operates in search as in reasoning