SYNTHESIS NOTE
Topics›RLVR›this note

Does RLVR actually expand what models can reason about?

Explores whether reinforcement learning from verifiable rewards teaches models genuinely new reasoning skills or simply makes existing capabilities more reliable. Pass@k analysis suggests the latter.

Synthesis note · 2026-02-22 · sourced from RLVR

The strongest empirical challenge to the "RL teaches reasoning" narrative comes from pass@k analysis. At small k (e.g., k=1), RLVR models outperform their base models — they produce correct answers more reliably on any given attempt. But as k increases, base models consistently surpass RLVR models across all benchmarks and model families. The reasoning paths that RLVR models generate are already present in the base model's sampling distribution.

This reframes what RLVR actually does. Rather than expanding the frontier of solvable problems, RLVR narrows the sampling distribution toward correct solutions that were already accessible. The model learns to find correct paths more efficiently, not to reason in fundamentally new ways. Manual inspection confirms: for most problems where RLVR models succeed, the base model can produce at least one correct chain-of-thought.

Six popular RLVR algorithms (including GRPO, PPO variants) perform similarly and all remain far from optimal in leveraging the base model's potential — they converge on similar subsets of the base model's capability space. This suggests the bottleneck is not algorithmic but structural: on-policy RL with verifiable rewards optimizes sampling, not capability.

The contrast with distillation is sharp. Distillation from a stronger teacher can transfer genuinely new reasoning patterns, expanding the student's reasoning scope beyond what the base model could sample. Since Does RL teach reasoning or just when to use it?, the RLVR finding fits: activation is not creation. But distillation is creation — it writes new patterns into the model's distribution.

The practical implication: if you need capabilities the base model doesn't have, distillation from a stronger model is the path. If the base model can already solve the problem (given enough samples), RLVR makes it reliable. These are different tools for different gaps.

Inquiring lines that read this note 124

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can AI systems achieve real improvement without external human feedback? How do reward signal properties affect model reasoning and safety? Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones? What explains the gap between benchmark scores and true reasoning capability? Which reinforcement learning modifications most improve dialogue quality in language models? How do agents learn to distinguish valuable feedback from noise? Does pretraining establish the ceiling for what reward learning can improve? Why does AI verification capability persistently exceed generation capability? How does policy entropy collapse limit scaling of reasoning-focused reinforcement learning? How can we reduce inherent biases in LLM-based evaluation judges? How does scaling reasoning capabilities affect models' appropriate abstention behavior? What prediction granularity best trains models to generate reliable reasoning? What makes process supervision effective for training complex reasoning models? What are the fundamental limits of prompting for language models? How does diversity prevent model convergence on superficial patterns? How do training data quality and composition affect downstream model performance? What limits recursive self-improvement in autonomous AI systems? Can external verification systems adequately replace learned reasoning in AI outputs? How does optimization for reward create emergent misalignment in language models? How do models learn from self-generated outputs without cascading failures? How do curriculum design and feedback approaches affect model learning? Can iterative DPO substitute for online RL in studying misalignment? How do AI systems determine and balance multiple competing objectives? What external process records should verify agent behavior and benchmark claims? Can AI research automation sustain progress through accelerating feedback loops? Does scaling reasoning capability create fundamental tradeoffs in control and reliability?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
22 direct connections · 167 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

rlvr does not expand reasoning capability boundaries beyond the base model — it improves sampling efficiency within existing boundaries