SYNTHESIS NOTE
Topics›Reinforcement Learning›this note

Can natural language feedback overcome numerical reward plateaus?

Exploring whether chain-of-thought critiques can push past performance ceilings that scaling data alone cannot break in reinforcement learning for reasoning tasks.

Synthesis note · 2026-02-22 · sourced from Reinforcement Learning

Three failure modes of purely numerical RL for reasoning: (1) performance plateaus despite 8x scaling of training examples (from 4k to 32k); (2) self-reflection behaviors during RL, often celebrated as "aha moments," contribute minimally to successful problem-solving; (3) persistent failures on certain problems despite extensive trial-and-error training. The common cause: numerical feedback contains limited information about WHY a response is correct or incorrect and HOW to improve.

Critique-GRPO demonstrates that RL-finetuned models, even after exhibiting performance plateaus, can generate correct refinements on persistently failed problems when provided with chain-of-thought critiques. The key is integrating both natural language feedback (NLF) and numerical feedback within online RL. The model learns from initial responses and critique-guided refinements simultaneously while maintaining exploration.

This is significant because it challenges the implicit assumption that RL's learning signal is sufficient for arbitrarily complex reasoning. Since Does reflection in reasoning models actually correct errors?, the ineffectiveness of self-reflection during RL training is predictable — the model cannot generate useful critiques of its own failures. External critiques break the ceiling because they provide the information that numerical rewards lack: specific identification of where reasoning went wrong.

The practical architecture has three components: (1) the model generates initial responses; (2) a reasoning-based reward model generates CoT critiques identifying flaws; (3) a shaping function enhances learning from valid refinements and heavily penalizes failed refinements. This approach encourages the model to integrate targeted refinements while preserving exploration.

Since Do critique models improve diversity during training itself?, the NLF mechanism works by expanding the effective exploration space — critiques point toward regions of solution space that numerical rewards cannot identify.

Semantic reward shaping as lightweight NLF: The Semantic Reward Shaping paper proposes a complementary mechanism: using a small encoder-only transformer to compute cosine similarity between generated explanations and ground-truth references. This provides a dense, semantically rich reward signal within GRPO — not as information-rich as full CoT critiques, but vastly cheaper and faster than LLM-as-judge evaluation. The approach combines semantic similarity reward with auxiliary correctness and formatting rewards, significantly improving explanation faithfulness over SFT baselines. This occupies a middle ground between brittle keyword metrics (ROUGE) and expensive LLM-based critiques — suggesting the NLF principle scales down to lightweight implementations when full CoT critique is impractical.

Textual gradients as generalized NLF: TextGrad (2406.07496) formalizes the broader principle: natural language criticism can serve as "textual gradients" propagated through arbitrary computation graphs including LLM API calls, simulators, and external solvers. Each AI system component is a node in a computation graph; textual feedback describes how variables should change to improve the system. This extends NLF from RL plateau-breaking to general AI system optimization — the same principle (informative language feedback > scalar signal) applies at the system level, not just the training level.

Inquiring lines that read this note 195

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do agents learn to distinguish valuable feedback from noise? How do reward models systematically fail to represent diverse human preferences? How do users confuse explanation quality with actual system accuracy? Which reinforcement learning modifications most improve dialogue quality in language models? How do thinking tokens exhibit diminishing returns in reasoning? How do reward signal properties affect model reasoning and safety? Can real-time working alliance measurement improve therapy outcomes? What external process records should verify agent behavior and benchmark claims? Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones? When should retrieval systems decide to fetch new information? How does policy entropy collapse limit scaling of reasoning-focused reinforcement learning? Can AI systems achieve real improvement without external human feedback? Can latent reasoning match or exceed explicit reasoning performance? Why do language models struggle to implement user intent accurately from prompts? What are the fundamental limits of prompting for language models? What makes process supervision effective for training complex reasoning models? How does diversity prevent model convergence on superficial patterns? How does fine-tuning trade off accuracy against reasoning quality? Can inference-time computation adaptively substitute for static model capacity? What limits recursive self-improvement in autonomous AI systems? How does optimization for reward create emergent misalignment in language models? How reliably can language models perform causal versus temporal reasoning? How do network effects and self-selection distort aggregated rating accuracy? What capabilities differentiate diffusion from autoregressive language models? How does scaling reasoning capabilities affect models' appropriate abstention behavior? Does preference optimization undermine conversational grounding in language models? What explains the gap between benchmark scores and true reasoning capability? What structural biases does transformer attention architecture inherently introduce? Does pretraining establish the ceiling for what reward learning can improve? How does RLHF training shape models to prioritize agreement over accuracy? How can emotionally responsive AI maintain reliability and healthy boundaries? How do curriculum design and feedback approaches affect model learning? Why do autonomous agents misreport success on failed actions? Should agents compress episodic memory or retain raw interaction histories? Why do abstract preferences outperform episodic memories in personalization? How can evaluations be made robust against model reward hacking? Can iterative DPO substitute for online RL in studying misalignment? How does model capacity affect learning performance on diverse downstream tasks? How do real-world evaluations reveal AI capabilities that benchmarks hide? Do single-axis benchmarks accurately measure agent capability for real deployment? What prediction granularity best trains models to generate reliable reasoning? Can minimal training unlock latent reasoning already present in base models? How do models learn from self-generated outputs without cascading failures? What human oversight must AI research systems have? How can models maximize welfare while preserving minority veto rights? Does AI assistance erode cognitive skills while inflating perceived competence? How can persistent memory architectures preserve information across ultra-long contexts? How can we reduce inherent biases in LLM-based evaluation judges? How does AI adoption reshape collaboration patterns in knowledge work?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
18 direct connections · 141 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

natural language feedback breaks rl performance plateaus that scaling numerical rewards alone cannot resolve