SYNTHESIS NOTE
Topics›Reinforcement Learning›this note

Does negative reinforcement alone outperform full reinforcement learning?

Can training with only penalty signals for wrong answers match or exceed full RL approaches? This challenges the conventional assumption that reward design requires both positive and negative signals.

Synthesis note · 2026-02-22 · sourced from Reinforcement Learning

Decomposing RL's learning signal into positive sample reinforcement (PSR) and negative sample reinforcement (NSR) reveals a surprising asymmetry. Training with only negative samples — penalizing incorrect responses without ever reinforcing correct ones — consistently improves performance over the base model across the entire Pass@k spectrum (k up to 256), often matching or surpassing full PPO and GRPO.

The mechanism is straightforward through gradient analysis: NSR works by suppressing incorrect generations and redistributing probability mass toward other plausible candidates, guided by the model's prior beliefs. It refines existing knowledge rather than introducing entirely new behaviors. This is because penalizing a wrong answer doesn't point toward any specific correct answer — it lets the model's own prior determine where the freed probability mass flows.

Positive-only reinforcement creates the opposite problem. It improves Pass@1 (the model gets better at its top-ranked answer) but degrades performance at higher k because it concentrates probability mass on rewarded trajectories, reducing diversity. Since Does policy entropy collapse limit reasoning performance in RL?, positive reinforcement actively contributes to the problem while negative reinforcement sidesteps it.

This reframes how we think about RL for reasoning. The conventional framing is that RL rewards correct behavior. But the evidence suggests that penalizing incorrect behavior may contribute more to performance than reinforcing correct behavior — especially when diversity matters. The model already contains good solutions in its prior; it just needs help avoiding the bad ones.

The practical implication is that reward design for reasoning RL may be over-engineered. If suppression alone gets you most of the way, the elaborate reward shaping and process supervision architectures may be solving a problem that's already largely solved by the base model's prior distribution.

Inquiring lines that read this note 85

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do agents learn to distinguish valuable feedback from noise? When do simpler collaborative filtering approaches outperform complex LLM recommenders? How do reward signal properties affect model reasoning and safety? Which reinforcement learning modifications most improve dialogue quality in language models? Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones? How do reward models systematically fail to represent diverse human preferences? What makes process supervision effective for training complex reasoning models? How do training data quality and composition affect downstream model performance? How do curriculum design and feedback approaches affect model learning? How does scaling reasoning capabilities affect models' appropriate abstention behavior? Can AI systems achieve real improvement without external human feedback? Does pretraining establish the ceiling for what reward learning can improve? Why do training associations persist despite contradictory contextual information? How does RLHF training shape models to prioritize agreement over accuracy? How does awareness of evaluation context influence model behavior? How does optimization for reward create emergent misalignment in language models? How do models learn from self-generated outputs without cascading failures?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 125 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

negative reinforcement alone matches or exceeds full rl by suppressing incorrect trajectories and redistributing probability mass