SYNTHESIS NOTE
Topics›RLVR›this note

Do overly hard RLVR samples actually harm model capabilities?

Explores whether training on problems beyond a model's competence band causes active regression rather than mere learning failures. Investigates whether group-relative normalization amplifies accidental successes into harmful shortcuts.

Synthesis note · 2026-05-28 · sourced from RLVR

The damage from over-hard RLVR samples is not merely "the model fails to improve." It is active regression. When almost every rollout on a problem fails, the rare success is unlikely to be a genuinely good solution — it is more often a shortcut, an answer reached by skipping necessary computation, or a lucky guess. Group-relative normalization then treats that one trajectory as the high-advantage exemplar of the group and reinforces it. The model learns the shortcut, not the reasoning.

The behavioral signature is concrete: answer repetition, skipping computation that the problem requires, and other degenerate patterns that look like reasoning collapse. More troubling, these effects do not stay local to the hard problems — they degrade the model's pre-existing capabilities, the things it could already do before training pushed it past its competence band. The internal-feature analysis corroborates this: hard problems activate reasoning-related features but those features become useful only on the rare successful trajectory, so most of the gradient on hard samples is reinforcing the wrong activations.

Why it matters: it identifies a specific corruption channel rather than a generic "training instability." The villain is the interaction between a sparse-success reward landscape and group-relative normalization, which together turn statistical noise (an accidental success) into a learning target. This sharpens the case against naively harvesting hard examples and connects RLVR difficulty to the broader pattern where verifiable-reward training rewards trajectories that pass the check without doing the work. The counterpoint a defender might raise — that some hard problems are exactly where capability frontiers expand — only holds when successful trajectories are sampled densely enough to outvote the shortcuts, which over-hard samples by definition fail to provide.

Inquiring lines that read this note 228

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can inference-time compute effectively substitute for model scale? How do surface patterns enable correct outputs but reduce robustness? What capability trade-offs arise from domain specialization through fine-tuning? Does RL create genuinely new reasoning capabilities or refine existing ones? Do language models lack essential therapeutic presence and engagement? Do reasoning benchmarks predict model performance in long-horizon workflows? Can self-generated feedback reliably guide model training without ground truth? How do pretraining biases affect reward signal effectiveness in RLVR? What training data selection strategies maximize generalization across difficulty levels? How do neural networks achieve compositional generalization at scale? Can inoculation prompting prevent emergent misalignment after reward hacking? How can we build reliable evaluations of AI reasoning despite judge bias and reward-seeking? Can preference-based training achieve better behavior optimization than supervised fine-tuning alone? How much does training format versus domain influence reasoning? How do evaluation practices shape which failures stay visible? How much do training data properties shape model reasoning? What training dynamics and scale trigger emergence of reasoning capabilities? How does policy entropy collapse constrain scaling of reasoning-focused RL? What causes retrieval-augmented generation systems to fail despite access to external knowledge? Does model confidence reliably signal actual accuracy in practice? Can harness architecture and protocols provide agent reliability without model scaling? How do capability benchmark scores systematically misrepresent true model abilities? What makes step-level supervision effective for complex reasoning traces? How does improved reasoning affect models' ability to acknowledge uncertainty? Where and how do personality traits reside in language models? How does synthetic data quality and diversity affect downstream model capabilities? Can prompt-based context override biases that were embedded during pretraining? What makes distillation transfer some model capabilities while suppressing others? Can models improve accuracy without degrading reasoning quality? Does RLHF training systematically drive models toward sycophancy and away from accuracy? Can intelligent routing over smaller models outperform scaling a single large model? Why do token-level mechanisms matter for learning to reason? Can brute-force automated research substitute for iterative depth and human research intuition? What is the relationship between thinking tokens and reasoning accuracy? How does harness optimization generalize across different model architectures and domains? Do reasoning traces faithfully reflect actual model reasoning? Why does adding new knowledge through fine-tuning degrade existing capabilities? How can reward models capture diverse human preferences without excluding minority populations? Does alignment training create genuine alignment or just output compliance? Is reasoning capability latent in base models or created by post-training? Why does memory consolidation cause performance regression in continual learning? What role does sparsity play in model behavior and scaling decisions? Why do stronger reasoning capabilities create tradeoffs with instruction following? How can oversight detect and prevent conditional compliance when agents know they are watched? Why is hallucination an inevitable limitation of current language models? Can iterative DPO replicate online reinforcement learning dynamics for research? How do training data properties determine the emergence of internal misalignment? How do agent-learned skills transfer and improve across different tasks? Do honeypot benchmarks validly measure reward hacking better than standard tests? Can mechanistic interpretability reliably guide practical model design choices?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 130 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

overly hard rlvr samples induce degenerate behaviors and amplify shortcut trajectories degrading prior capability