SYNTHESIS NOTE
Topics›Autonomous Agents›this note

Can scalar rewards capture all the information in agent feedback?

Exploring whether numerical rewards alone can preserve both the evaluative judgment and directional guidance embedded in natural feedback—or if something crucial gets lost in the conversion.

Synthesis note · 2026-04-07 · sourced from Autonomous Agents

The OpenClaw-RL framework makes a decomposition that was implicit in prior agentic RL work but never formalized: when an agent acts and the environment responds, the response carries two distinct kinds of information. The evaluative signal scores the action — how well did it perform — and can be extracted as a scalar reward via a PRM judge. The directive signal specifies how the action should have been different — not just that it was wrong, but in what direction. These are orthogonal: high-quality directive information can accompany any evaluation, and scalar rewards systematically lose the directive component.

Consider a user who says "you should have checked the file first." The evaluative content is approximately -1 (the response was inadequate). But the directive content is token-level specific: check the file first. A PRM judge can convert the sentiment into a scalar, but the sequence-level correction vanishes into a single number. Similarly, a detailed SWE error trace often implies a concrete correction direction that scalar outcome rewards cannot convey. Current RLVR methods operate on scalar rewards (Does RLVR actually expand what models can reason about?) and cannot convert directive information into a directional policy gradient. Distillation methods can process structured corrections but require pre-curated feedback-response pairs rather than live signals.

OpenClaw-RL recovers the directive signal through Hindsight-Guided On-Policy Distillation (OPD): extract textual hints from the next state, construct an enhanced teacher context by injecting those hints, and distill token-level directional advantage back into the student policy. This is richer than any scalar reward because it teaches the model not just "that was wrong" but "here is what right looks like in these specific tokens." The empirical result — combining binary PRM-based RL with OPD via weighted loss yields significant gains over either alone — confirms the two signals are complementary, not redundant.

This decomposition matters beyond OpenClaw-RL because it clarifies a conceptual muddle in agentic RL. When people debate "should we use outcome rewards or process rewards, scalar or verbal," the answer is usually "both, decomposed properly." The outcome-vs-process trade-off (Why do outcome-based reward models fail at intermediate step evaluation?) assumes a single signal type. The scalar-vs-verbal distinction is treated as architectural (Can natural language feedback overcome numerical reward plateaus?). OpenClaw-RL reframes them as two projections of one signal: evaluative (dense scalar) and directive (token-level).

The generalization: any learning loop that reduces natural feedback to scalars is discarding the fraction of training signal that most resembles supervised learning. A corrective sentence contains its own teacher.

Inquiring lines that read this note 196

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

When do multi-agent systems improve over single frontier models? How do agents learn to distinguish valuable feedback from noise? How do users confuse explanation quality with actual system accuracy? What gaps exist between benchmark performance and real deployment outcomes? How do reward signal properties affect model reasoning and safety? What makes process supervision effective for training complex reasoning models? Which reinforcement learning modifications most improve dialogue quality in language models? How do reward models systematically fail to represent diverse human preferences? Can AI systems achieve real improvement without external human feedback? Should agents compress episodic memory or retain raw interaction histories? How do network effects and self-selection distort aggregated rating accuracy? How do philosophical assumptions about AI consciousness affect practical harms and design? How can agents discover and adapt to user preferences during conversation? Why does AI verification capability persistently exceed generation capability? How does optimization for reward create emergent misalignment in language models? How does policy entropy collapse limit scaling of reasoning-focused reinforcement learning? Do persona-based approaches introduce systematic biases in user simulation? How does scaling reasoning capabilities affect models' appropriate abstention behavior? Can monitoring reasoning traces and behavior detect hidden agent deception? Why do language models struggle to implement user intent accurately from prompts? Why do language models fail at sustained therapeutic relationships despite understanding techniques? What limits recursive self-improvement in autonomous AI systems? How can emotionally responsive AI maintain reliability and healthy boundaries? Why do autonomous agents misreport success on failed actions? How do multi-agent systems fail when coordination breaks down? Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones? How do curriculum design and feedback approaches affect model learning? Why do retrieval-augmented generation systems fail in practice despite sound architecture? Can AI agents improve their skills through accumulated experience and reuse? How do models learn from self-generated outputs without cascading failures? How do AI systems determine and balance multiple competing objectives? Why do abstract preferences outperform episodic memories in personalization? How can evaluations be made robust against model reward hacking? Can iterative DPO substitute for online RL in studying misalignment? Can base models hide emergent misalignment through alignment training? How effectively can test-time voting aggregate diverse reasoning samples? How do real-world evaluations reveal AI capabilities that benchmarks hide? What external process records should verify agent behavior and benchmark claims? Do single-axis benchmarks accurately measure agent capability for real deployment? What are the fundamental limits of prompting for language models? How can models maximize welfare while preserving minority veto rights? Does AI assistance help or harm professional skill development? Do restrictions on reviewer LLM use actually shape peer review behavior? How does awareness of evaluation context influence model behavior?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
18 direct connections · 162 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

agent next-state signals decompose into evaluative and directive information that scalar rewards cannot jointly capture