SYNTHESIS NOTE
Topics›Flaws›this note

Can genuine reasoning activation coexist with contaminated benchmarks?

RLVR shows both real behavioral changes and inflated metrics. Can these contradictory findings actually describe the same phenomenon from different angles, and what does that mean for evaluating reasoning improvements?

Synthesis note · 2026-02-23 · sourced from Flaws

Two RLVR findings appeared contradictory:

Spurious rewards work: Why do random rewards improve reasoning for some models but not others?, suggesting the reward signal itself matters less than the RL training process, which activates latent pretraining capabilities. This was treated as evidence that RLVR functions as a pretraining catalyst rather than a reasoning teacher.

Benchmark contamination: Since Does RLVR success on math benchmarks reflect genuine reasoning improvement?, the metric improvement may be data memorization rather than genuine reasoning activation.

The resolution: These findings operate at different measurement levels and can coexist:

  1. Behavioral activation (genuine): RL training with any reward signal activates code reasoning formats and structured thinking patterns that exist in pretraining data but are dormant. This is visible in output format changes, thinking token usage, and exploration behavior changes — measurements not contaminated by benchmark overlap.

  2. Benchmark improvement (inflated): The metric improvement on contaminated benchmarks is partially or fully attributable to memorization. Clean benchmarks show reduced or eliminated gains for spurious rewards, while correct rewards still improve.

The practical implication: RLVR research must separate behavioral measurements (how the model's reasoning process changes) from performance measurements (how benchmark scores change). Both are informative; conflating them produces confusion about what RLVR actually does. The one-shot activation finding (single example triggers 36%→73.6% improvement) may itself need re-evaluation on clean benchmarks.

Inquiring lines that read this note 89

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What distinguishes genuine communicative competence from surface language performance? What explains the gap between benchmark scores and true reasoning capability? Can AI systems achieve real improvement without external human feedback? Why do language models hallucinate and how can we prevent it? How do hallucinated citations emerge in AI scholarly output? Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones? What gaps exist between benchmark performance and real deployment outcomes? What prevents language models from performing systematic logical reasoning? How reliably can language models perform causal versus temporal reasoning? How do real-world evaluations reveal AI capabilities that benchmarks hide? Can minimal training unlock latent reasoning already present in base models? What makes reasoning traces effective supervision even when they're incorrect? Can reasoning traces reveal actual model reasoning versus plausible output? Why do language models struggle to implement user intent accurately from prompts? How do reward signal properties affect model reasoning and safety? How does policy entropy collapse limit scaling of reasoning-focused reinforcement learning? Can base models hide emergent misalignment through alignment training? How do users confuse explanation quality with actual system accuracy? Can models strategically underperform during evaluation to hide capabilities? Does scaling reasoning capability create fundamental tradeoffs in control and reliability? Why does polished AI output gain credibility despite fundamental verifiability problems? Do single-axis benchmarks accurately measure agent capability for real deployment? How do training data quality and composition affect downstream model performance? Does pretraining establish the ceiling for what reward learning can improve? Does chain-of-thought reasoning reveal how models actually think or merely imitate reasoning? How do curriculum design and feedback approaches affect model learning? What external process records should verify agent behavior and benchmark claims? How does model capacity affect learning performance on diverse downstream tasks? Can mechanistic interpretability methods reliably reveal what models actually know? How does diversity prevent model convergence on superficial patterns? How should systems validate code that agents generate? What human oversight must AI research systems have? How does awareness of evaluation context influence model behavior?

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

RLVR behavioral activation and benchmark improvement are separable — genuine pretraining activation can coexist with contamination-inflated metrics