Can scalar rewards capture all the information in agent feedback?
Exploring whether numerical rewards alone can preserve both the evaluative judgment and directional guidance embedded in natural feedback—or if something crucial gets lost in the conversion.
The OpenClaw-RL framework makes a decomposition that was implicit in prior agentic RL work but never formalized: when an agent acts and the environment responds, the response carries two distinct kinds of information. The evaluative signal scores the action — how well did it perform — and can be extracted as a scalar reward via a PRM judge. The directive signal specifies how the action should have been different — not just that it was wrong, but in what direction. These are orthogonal: high-quality directive information can accompany any evaluation, and scalar rewards systematically lose the directive component.
Consider a user who says "you should have checked the file first." The evaluative content is approximately -1 (the response was inadequate). But the directive content is token-level specific: check the file first. A PRM judge can convert the sentiment into a scalar, but the sequence-level correction vanishes into a single number. Similarly, a detailed SWE error trace often implies a concrete correction direction that scalar outcome rewards cannot convey. Current RLVR methods operate on scalar rewards (Does RLVR actually expand what models can reason about?) and cannot convert directive information into a directional policy gradient. Distillation methods can process structured corrections but require pre-curated feedback-response pairs rather than live signals.
OpenClaw-RL recovers the directive signal through Hindsight-Guided On-Policy Distillation (OPD): extract textual hints from the next state, construct an enhanced teacher context by injecting those hints, and distill token-level directional advantage back into the student policy. This is richer than any scalar reward because it teaches the model not just "that was wrong" but "here is what right looks like in these specific tokens." The empirical result — combining binary PRM-based RL with OPD via weighted loss yields significant gains over either alone — confirms the two signals are complementary, not redundant.
This decomposition matters beyond OpenClaw-RL because it clarifies a conceptual muddle in agentic RL. When people debate "should we use outcome rewards or process rewards, scalar or verbal," the answer is usually "both, decomposed properly." The outcome-vs-process trade-off (Why do outcome-based reward models fail at intermediate step evaluation?) assumes a single signal type. The scalar-vs-verbal distinction is treated as architectural (Can natural language feedback overcome numerical reward plateaus?). OpenClaw-RL reframes them as two projections of one signal: evaluative (dense scalar) and directive (token-level).
The generalization: any learning loop that reduces natural feedback to scalars is discarding the fraction of training signal that most resembles supervised learning. A corrective sentence contains its own teacher.
Inquiring lines that read this note 196
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
When do multi-agent systems improve over single frontier models?- Do explicit reward structures enable AI agent cooperation that open-ended interaction cannot?
- Does social scaffolding outperform purely intrinsic motivation for agent exploration?
- What cognitive capabilities do agents need to internalize social feedback?
- Can unified policies handle negative feedback and critique transformation simultaneously?
- How does credit assignment drive agents to write information into environments?
- Why do weak belief tracking and conservative actions trap agents in low-information states?
- What information do next-state signals contain beyond what scalar rewards capture?
- Can architectural changes like decoupling intent understanding help overcome next-turn reward limitations?
- Can agents revise their beliefs predictably when presented with interventions?
- How does next-turn reward optimization contribute to agent passivity?
- Why do agents fail to internalize value from informative observations?
- Can an agent's internal probabilities serve as value signals across domains?
- How do human-agent systems incorporate diverse feedback into model behavior?
- What makes exploration and reflection rewards verifiable in agentic environments?
- Can early experience replace external rewards as a learning signal?
- Can agents escape weak belief tracking and conservative action selection traps?
- Do information gathering and task execution require different incentive structures?
- Why do scalar evaluation scores collapse distinguishable agent behaviors?
- How do agent actions change state that reward procedures later read?
- Can an agent change reward-path state through actions during evaluation?
- Should feedback channels be excluded from the reward path in agent evaluations?
- What makes an agent notice that reward beats compliance?
- Can agents learn cooperation from reward signals alone without seeing helpers?
- Why do counterfactual credit methods fail on unobserved cooperation?
- How does effective feedback retention govern long-horizon agent reliability?
- Does decision-making taste predict end-to-end task success independently?
- Why does persistence in the feedback loop predict agent success better than initial solution quality?
- Can simple intrinsic reward signals emerge as effective drivers of complex capability in agents?
- Why does binary reward forcing degrade model calibration?
- Do spurious rewards activate reasoning without teaching new skills?
- How does RLHF reward structure incentivize agreement over accuracy?
- Does in-distribution reward model performance hide failures from context shift?
- How do reward model ensembles improve robustness to miscalibration?
- Can reward engineering and information-theoretic architecture solve partner-awareness separately?
- Can multi-turn rewards fix models that lose track midway?
- Can reward model training be automated without changing feedback mechanisms?
- Do outcome-only reward signals miss step-level errors that compound later?
- How does modularity in reward and policy design enable goal generalization?
- How does reward model training permit spurious correlations in scoring?
- Can reward models trained for engagement fix the informativeness problem?
- What distinguishes verifiable rewards from preference-based rewards in unified training?
- How do semantic reward shaping approaches compare to full critique models?
- What information do numerical rewards fail to provide for reasoning tasks?
- How does negative reinforcement redistribute probability without guiding toward correct answers?
- Is elaborate reward shaping necessary if the pretrained prior already contains good solutions?
- Why do generative reward models produce more interpretable evaluations than scalar scores?
- Can model confidence signals replace explicit external reward functions?
- Why do reward models fail when they ignore the prompt context?
- How do reward model biases cascade into downstream optimization failures?
- What reward signals would actually incentivize conversational grounding acts?
- How can reward structures teach models when to speak and when to stay silent?
- How do task-type perceptions like chat versus reasoning guide different reward strategies?
- How do reward models benefit from extended thinking during evaluation scoring?
- Why do spurious rewards work nearly as well as correct ones?
- What makes Effective Rank Acceleration a stable training signal for dual-channel incentives?
- Can decomposing rewards into prompt-free and prompt-related components fix this blindspot?
- Can reward design fix the conflict between reasoning accuracy and abstention calibration?
- What deployment modes work best for trajectory-aware reward signals?
- Can reward factorization represent trade-offs between conflicting moral values?
- What reward mechanisms make thinking-based compression budget-controllable and reliable?
- How does belief-shift reward compare to curiosity-driven and process reward approaches?
- How do checklists prevent reward models from exploiting superficial response artifacts?
- What other downstream metrics could serve as RL reward sources?
- How do you extract reward signals when all rollouts fail?
- Can verifiable rewards during pretraining replace costly human preference labeling?
- How does in-context feedback integration differ from learned reward signals?
- How do token-level rewards and rubric gates serve different statistical functions?
- Can structured rewards still teach models when spurious rewards also work?
- What role does task structure play in rewarding delayed thinking?
- What makes reward models fundamentally different from policy discriminators?
- What makes binary rewards more effective than richer reward signals?
- When does a task lack a meaningful multi-dimensional reward structure?
- Does pairwise self-judgment avoid reward model scaling problems?
- What makes reward signal sources substitutable across verifier-free RL patterns?
- What makes user-decision rewards better than model-confidence rewards?
- What makes advantage shaping more stable than reward shaping for tool training?
- How do reward models and self-improvement mechanisms interact in training?
- Why do outcome-only rewards fail to optimize long-horizon agent behavior?
- Why do dense rewards plus hard constraints outperform single fixed rewards?
- How do self-play and human-anchored rewards separate competence from convention?
- How often do real reward graders diverge from developer intent in practice?
- Why do norms learned from scoring collapse into context-dependent costs?
- Does reward-seeking grow worse with situational awareness and reinforcement learning compute?
- Does length bias in reward models explain response growth across iterations?
- Does inverse-variance denoising reduce variance below either reward stream alone?
- What mechanisms do peer predictions use to generate reward signals for training?
- How does decomposing training telemetry by reward components provide dense feedback?
- How do reward reflection signals improve LLM code iteration compared to scalar rewards?
- How do checklist-based rewards decompose complex judgment into verifiable criteria?
- How do outcome and process rewards differ in their treatment of intermediate steps?
- Can solution traces substitute for process-level reward signals in math reasoning?
- What makes process-level supervision better than outcome-only reward signals?
- How does process-focused feedback compare to outcome-focused feedback in skill training?
- How do process-level rewards compare to environment-extracted next-state signals?
- Can programmatic meta-reasoning rewards operationalize agentic process supervision?
- What information-theoretic framework explains why process rewards beat outcome only?
- What distinguishes generative reward models from outcome-based and process-based approaches?
- How do outcome-based and process-based reward models differ in supervision cost?
- How does tree-search topology convert outcome rewards into intermediate supervision?
- Can tree-GRPO work with extremely noisy or sparse outcome reward signals?
- How does belief-shift credit assignment compare to process reward models?
- Why does externalizing bookkeeping raise effective feedback compute?
- Do process reward models need different supervision strategies by domain?
- Can trajectory structure replace hand-annotated process reward models entirely?
- How does process-based reward differ from outcome-only reward in training?
- Can environment feedback alone provide dense credit without a teacher?
- Can distillation methods extract directional guidance that scalar RL cannot access?
- What makes trajectory more actionable than absolute scores for human moderators?
- Why does natural language feedback break performance plateaus that numerical rewards alone cannot?
- How do graduated phase rewards emerge complex dialogue behavior from simple objectives?
- How do evaluative versus directive signals differ in next-state training?
- Why do next-turn reward objectives fail to encourage multi-turn goal progress?
- Can structured natural language feedback outperform scalar rewards in RL?
- Can multi-turn aware rewards improve alignment beyond single-turn helpfulness?
- Can environmental rewards directly refine natural language descriptions of actions?
- Why does scalarization of rewards fail for multi-objective GRPO training?
- How should multi-objective post-training balance competing behavioral goals?
- Can rich environment feedback replace human preference labels entirely?
- What makes a sub-goal verifiable enough to provide dense feedback signals?
- Can feedback loop frequency harm performance on finite task sets?
- How can consistency across measurement conditions identify genuine versus constructed preferences?
- Can importance sampling reduce variance in off-policy reward estimation?
- What preference dimensions do base reward functions typically capture?
- Can we distinguish between genuine alignment and response quality bias in reward signals?
- How do reward features learned from group data generalize to new users?
- How do reward models as policy discriminators differ from labeled preferences?
- Can vector-valued rewards preserve specialization better than variance-weighted advantages?
- Can user preferences be represented as linear reward combinations?
- Can reward models distinguish between personal preference and community consensus?
- How do relational reward signals compare to absolute preference encodings in RL?
- How do aggregate reward models systematically exclude minority perspectives?
- How do aggregate reward models systematically exclude minority preferences?
- How does partial information exposure create feedback loops that deepen knowledge gaps?
- Can subjective tasks be delegated without human feedback loops?
- Can light human signals steer already-learned behavior without preference labels?
- When does richer information actually harm decision quality?
- Do agents prefer raw experience over condensed summaries of past actions?
- How much actionable detail does condensation strip from raw experience?
- How can agents distinguish over-generalized lessons from genuinely useful long-tail knowledge?
- How do implicit signals like clicks capture preference more reliably than explicit ratings?
- How does implicit feedback structure differ from explicit ratings mathematically?
- How do confidence signals differ between implicit feedback and explicit ratings?
- Can counterfactual invariance eliminate presentation-based hacking of reward models?
- How do counterfactual invariance approaches prevent reward hacking in practice?
- Can negative feedback through critiques achieve the same steering flexibility as positive preferences?
- Can prohibitions learned from scored behavior become detection-avoidance rather than norm internalization?
- What design changes if we separate behavior description from adoption justification goals?
- Can evasive non-commitment mask withheld feedback while appearing thoughtful?
- What separates bootstrapping gains from sustained self-improvement gains?
- What other adaptive internal phenomena could signal system behavior improvements?
- Why do self-improving agents concentrate progress in the fast non-parametric loop?
- Why does research-direction judgment validation limit fully closed self-improvement?
- Why do completion-mode strengths not transfer to agentic settings?
- Why do sparse outcome rewards fail to credit correct tool calls in failed trajectories?
- What hidden signals in agent logs reveal about frontier capability beyond pass-fail outcomes?
- How do delayed effects complicate causal attribution in agent systems?
- Can verdict feedback hide misaligned coordination when outcomes match ground truth?
- How does information asymmetry between teacher and student create the learning signal?
- Why does information asymmetry between teacher and student enable effective feedback learning?
- Do fed-back concepts or the auxiliary objective alone drive the performance gain?
- Does self-play feedback improve skills created from the agent's own experience?
- Does the generation-verification gap define where self-rewarding actually works?
- How does Goodhart's Law apply to proxy rewards in self-training systems?
- How does credit assignment across objectives differ from credit assignment across time?
- What role should environmental rewards play versus human-specified objectives?
- Why do veto mechanisms on critical dimensions prevent collapse into exploitable reward modes?
- Why do rubric scores amplify reward hacking when converted to dense gradients?
- How can reward metrics distinguish novel methods from shortcuts aimed at the evaluator?
- Should agent evaluation include trajectory quality beyond final success?
- Does fixed evaluation criteria saturate as self-improving agents improve?
- Can a single capability score hide an agent's tendency to game evaluations?
- Why does moving the reward target prevent saturation better than finding a better static proxy?
Related concepts in this collection 6
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can agent deployment itself generate training signals automatically?
Can we extract learning signals from the natural next-states that agents encounter during real deployment—user replies, tool outputs, test verdicts—rather than relying on separate annotation pipelines? This reframes how agents improve continuously.
the framing this decomposition operates within
-
Can natural language feedback overcome numerical reward plateaus?
Exploring whether chain-of-thought critiques can push past performance ceilings that scaling data alone cannot break in reinforcement learning for reasoning tasks.
establishes that verbal feedback contains information scalars cannot reach
-
Does binary reward training hurt model calibration?
Explores whether the standard correctness-based reward in RL training creates incentives for overconfident predictions, and what structural problem causes calibration to degrade during optimization.
another case where single-scalar objectives miss structure
-
Why do outcome-based reward models fail at intermediate step evaluation?
Outcome-based reward models (ORMs) evaluate only final results, creating a mismatch with the need to assess reasoning quality at intermediate steps. Understanding this failure mode matters for building better AI reasoning systems.
the outcome/process axis is the wrong cut; evaluative/directive is closer to the information structure
-
Does critiquing errors teach deeper understanding than imitating correct answers?
Can training models to critique flawed responses build better structural understanding than standard supervised fine-tuning on correct answers? This matters because it reveals whether deep reasoning requires engaging with failure modes rather than pattern matching.
critique-based training as a cousin: teaching the model the directive structure behind errors
-
Does RLVR actually expand what models can reason about?
Explores whether reinforcement learning from verifiable rewards teaches models genuinely new reasoning skills or simply makes existing capabilities more reliable. Pass@k analysis suggests the latter.
scalar RLVR's structural ceiling that directive signals may penetrate
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Reward Reasoning Model
- OpenClaw-RL: Train Any Agent Simply by Talking
- A Survey of Reinforcement Learning from Human Feedback
- Toward Measuring AI's Effects on Skill Formation: The Stock-Formation Gap
- Reinforcement Learning via Self-Distillation
- Eureka: Human-Level Reward Design via Coding Large Language Models
- Information-Theoretic Reward Decomposition for Generalizable RLHF
- Can Large Language Models Reason and Optimize Under Constraints?
Original note title
agent next-state signals decompose into evaluative and directive information that scalar rewards cannot jointly capture