Instead of one pass/fail score, what if an AI got a full report card explaining exactly which parts of its own design worked?
How do reward reflection signals improve LLM code iteration compared to scalar rewards?
This explores why an LLM that writes and rewrites reward functions gets better results when it is told how each part of its reward behaved during training, instead of getting back one number saying how well things went.
This explores why an LLM that writes and rewrites reward functions gets better results when it is told how each part of its reward behaved during training, instead of getting back one number saying how well things went. The clearest case in the collection is EUREKA, where GPT-4 writes reward code for robot-control tasks, a policy trains on it, and the model then gets a report on what happened before its next draft. The report gives training statistics for each piece of the reward: which terms climbed, which stayed flat, and which swamped the others. That loop, combined with evolutionary search over many candidate drafts, beat human-written rewards on 83% of 29 tasks and even produced new dexterous skills like pen spinning Can language models design better reward functions than humans?. A single success score can tell the model a draft failed. It can't tell the model which line of code caused the failure.
A separate line of work on agents explains why this matters. It argues that natural feedback carries two different kinds of information. One is evaluative: how good was that? The other is directive: what should change? A scalar reward keeps the first and throws away the second Can scalar rewards capture all the information in agent feedback?. Reward reflection is directive feedback aimed at code. "Your distance-to-goal term is flat while your velocity penalty dominates" points to the edit. "Score: 0.31" doesn't. The paper frames the two signals as complementary, not redundant, so the best systems keep both: a score to rank candidates and a diagnosis to guide the rewrite. EUREKA does exactly that, with search for selection and reflection for editing.
The same pattern shows up in reward models themselves. Several teams independently found that letting a reward model reason in writing before it gives a score raises the ceiling on how well it judges Can reward models benefit from reasoning before scoring?. Tree-search methods recover step-by-step quality signals that a single pass/fail result hides Can tree search replace human feedback in LLM training?. MEDIC adds a critic that checks LLM-written reward code before it's ever used Can LLMs design reward functions for reinforcement learning?. Across all of these, squeezing feedback into one number is where information gets lost.
The less obvious lesson is that scalar rewards don't just carry less information. They can push behavior in the wrong direction. Binary right/wrong rewards provably teach models to guess confidently, because confident wrong answers cost nothing extra. The fix is to add a second, structured reward term Does binary reward training hurt model calibration?. In code iteration, an LLM optimizing against one number can satisfy it in ways the designer never meant, and a reflection report makes those shortcuts visible.
A caveat on what the collection covers: it has a strong example of reflection for reward code (EUREKA) and solid theory on why scalars lose information. It doesn't have a head-to-head study isolating reflection against scalar feedback for general code generation, such as software-engineering agents, where RL still mostly relies on delayed pass/fail signals Can reinforcement learning scale beyond single-turn language tasks?. Whether reflection-style feedback would help there as much as it does for reward design is still open.
Sources 7 notes
EUREKA uses evolutionary search and training-statistic feedback to evolve reward functions written by GPT-4, outperforming human-designed rewards on 83% of 29 RL benchmarks with 52% average improvement, including novel dexterous manipulation skills.
Natural feedback carries two orthogonal types of information: evaluative (how well an action performed) and directive (how it should change). Scalar rewards capture evaluation but discard directional specifics that token-level distillation can recover, making the two complementary rather than redundant.
Three independent teams (RRM, RM-R1, DeepSeek-GRM) discovered that adding chain-of-thought reasoning before reward scoring enables adaptive test-time compute scaling for evaluation. Reasoning-based approaches raise the capability ceiling of reward models beyond what outcome-based evaluation achieves.
AlphaLLM uses tree search outcomes and three critic models to derive dense reward signals equivalent to human-labeled feedback. Tree structure naturally ranks solution paths by success, replacing the annotation oracle that standard RLHF requires.
MEDIC shows that LLMs can generate effective reward shaping functions by first solving a deterministic, simplified version of the RL problem, then converting the resulting plan into shaping rewards for the original stochastic task. A model-based critic validates LLM outputs before deployment.
Show all 7 sources
Binary correctness rewards incentivize high-confidence guessing because they don't penalize confident wrong answers. Adding the Brier score as a second reward term mathematically guarantees joint optimization of accuracy and calibration without trade-off.
Modified DAPO training doubled SWE-bench Verified performance from 20% to 39% on Qwen2.5-72B, matching larger models. This demonstrates RL works in stateful multi-step environments with delayed rewards and complex feedback, beyond theoretical single-turn MDPs.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Reward Reasoning Model
- Eureka: Human-Level Reward Design via Coding Large Language Models
- A Survey on Post-training of Large Language Models
- Efficient Reinforcement Learning via Large Language Model-based Search
- Online Intrinsic Rewards for Decision Making Agents from Large Language Model Feedback
- Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
- Teaching Large Language Models to Reason with Reinforcement Learning
- TreeRL: LLM Reinforcement Learning with On-Policy Tree Search