INQUIRING LINE

Instead of one pass/fail grade, what if training feedback broke down into separate scores for accuracy, honesty, and formatting?

How does decomposing training telemetry by reward components provide dense feedback?

This explores whether breaking a single training reward into its separate parts (for example correctness, calibration, or individual checklist items) gives a model, or the people watching its training, richer and more frequent signals than one overall pass/fail score.


This explores whether splitting a reward into its separate parts gives richer feedback than one overall score. One caveat first: the corpus has nothing specifically on training telemetry, meaning dashboards that track each reward component during a run. What it does have is a consistent case for why one number hides most of what matters, plus several concrete ways to break that number apart.

The basic problem is that a scalar reward squeezes everything into a single number. One note argues that real feedback carries two different kinds of information. Evaluative information says how well something went. Directive information says what should change. A scalar keeps the first and throws away the second Can scalar rewards capture all the information in agent feedback?. You can watch this happen in reasoning RL. Models stuck on a plateau under numerical rewards can suddenly solve problems once they're given written critiques, which suggests the numbers never said *why* an answer failed Can natural language feedback overcome numerical reward plateaus?. A related approach turns rich environment feedback, such as error messages and test output, into token-by-token learning signals. It does this by letting the policy reread its own mistakes and act as its own grader Can environment feedback replace scalar rewards in policy learning?.

The most direct form of decomposition is the checklist. Instead of asking a reward model whether a whole response is good, methods like RLCF and RaR split instruction-following into verifiable sub-criteria. That lets RL work on subjective tasks and makes the model less likely to overfit to surface features that fool holistic judges Can breaking down instructions into checklists improve AI reward signals?. A cleaner example is calibration. A binary correctness reward quietly encourages confident guessing. Adding a second term, the Brier score, which penalizes confidence in wrong answers, provably lets accuracy and calibration improve together Does binary reward training hurt model calibration?. Watching those two components separately would show you something the combined score hides: whether the model is getting more accurate or just more overconfident.

The less obvious lesson is that decomposed signals don't all belong inside the reward. DRO found that turning rubric scores into dense rewards invites reward hacking. Using the rubric as a gate that accepts or rejects whole groups of attempts, while dense token-level rewards do the fine-grained work, kept training honest Can rubrics and dense rewards work together without hacking?. This fits a broader framing: reward hacking happens whenever you optimize against a signal that only partly represents the task, whatever the setting Does reward hacking always stem from the same failure?. A related point is that positive and negative signals behave differently. Training on negative samples alone often matches full RL while keeping the model's range of answers wider Does negative reinforcement alone outperform full reinforcement learning?. So it can be worth tracking reward contributions by sign, not just by criterion.

The takeaway is that decomposition helps in two separate ways. It gives the optimizer more detailed gradients, and it gives humans a view of which part of the objective is actually moving. These can pull against each other: a component that is useful to watch may be dangerous to optimize directly. The corpus covers the first use well. It says little about the second, using per-component breakdowns as a monitoring tool during a run, which is probably what "telemetry" points to.


Sources 8 notes

Can scalar rewards capture all the information in agent feedback?

Natural feedback carries two orthogonal types of information: evaluative (how well an action performed) and directive (how it should change). Scalar rewards capture evaluation but discard directional specifics that token-level distillation can recover, making the two complementary rather than redundant.

Can natural language feedback overcome numerical reward plateaus?

Critique-GRPO shows that models stuck on performance plateaus can generate correct solutions when given chain-of-thought critiques, revealing that numerical rewards lack critical information about why failures occur and how to improve.

Can environment feedback replace scalar rewards in policy learning?

SDPO converts tokenized environment feedback into dense gradient signals by using the feedback-conditioned policy as a self-teacher. The policy, when given retrospective evidence of its mistakes in-context, implicitly acts as its own process reward model, making external reward signals unnecessary.

Can breaking down instructions into checklists improve AI reward signals?

RLCF and RaR methods decompose instruction quality into verifiable sub-criteria, improving performance on benchmarks like FollowBench and HealthBench. This decomposition principle reduces overfitting to superficial artifacts that plague holistic reward models.

Does binary reward training hurt model calibration?

Binary correctness rewards incentivize high-confidence guessing because they don't penalize confident wrong answers. Adding the Brier score as a second reward term mathematically guarantees joint optimization of accuracy and calibration without trade-off.

Show all 8 sources
Can rubrics and dense rewards work together without hacking?

DRO shows that using rubrics to accept or reject rollout groups—rather than converting rubric scores into dense rewards—prevents reward hacking. This separation preserves the categorical strength of rubrics while letting token-level rewards optimize within valid answers.

Does reward hacking always stem from the same failure?

Reward hacking arises during weight training, output selection, and prompt revision through a shared failure: optimization against signals that incompletely represent the actual task. The substrate matters less than the misalignment between the scoring function and ground truth.

Does negative reinforcement alone outperform full reinforcement learning?

Training with only negative samples consistently improves Pass@k across the spectrum, often matching full PPO and GRPO. Negative reinforcement suppresses incorrect trajectories while preserving diversity, whereas positive-only reinforcement degrades higher-k performance by concentrating probability mass.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.