Does binary reward training hurt model calibration?
Explores whether the standard correctness-based reward in RL training creates incentives for overconfident predictions, and what structural problem causes calibration to degrade during optimization.
Binary correctness reward is the dominant approach to RL training for reasoning: correct answer earns 1, incorrect earns 0. RLCR identifies its structural flaw: a binary reward does not penalize high-confidence wrong answers. The reward for "correct with 99% confidence" equals the reward for "correct with 51% confidence." Therefore the model has no incentive to match its expressed confidence to its actual accuracy — high-confidence guessing is rational if it succeeds often enough.
The consequence is calibration degradation: models become more confident over the training run but not proportionally more accurate. On out-of-domain problems, where accuracy doesn't keep pace with confidence, the degradation produces higher rates of confident incorrect answers — what the paper frames as increased hallucination frequency.
The mathematical fix: add the Brier score as a second reward term alongside binary correctness. The Brier score is a proper scoring rule — it is uniquely maximized when predicted probabilities exactly match true outcome probabilities. The composite reward RLCR is therefore provably maximized only when the model (1) outputs the most likely correct answer AND (2) expresses a calibrated confidence estimate. The proof holds for any bounded proper scoring rule as the calibration term.
A surprising negative: the log-likelihood loss, also a proper scoring rule, does NOT have this property when combined with binary correctness reward — it can incentivize incorrect answers under specific confidence profiles. The bounded property of the Brier score is what enables the joint optimization guarantee.
Empirically: across diverse datasets, RLCR substantially improves calibration on both in-domain and out-of-domain evaluations with no accuracy cost. Standard RL hurts calibration; RLCR improves it.
The RLSF (RL with Self-Confidence) framework provides a complementary approach: instead of adding an external calibration term, it uses the model's own verbalized confidence as an intrinsic reward signal. The model generates a confidence estimate alongside its answer, and the reward combines correctness with confidence calibration. This is architecturally simpler than RLCR's Brier score approach but relies on the model's ability to self-assess — which Does reflection in reasoning models actually correct errors? suggests may be unreliable. RLCR's mathematical guarantee may be more robust than RLSF's empirical approach.
Two complementary robustness approaches from the reward hacking literature:
Bayesian Reward Model Ensembles (BRME) — train a multi-head reward model where each head outputs mean and standard deviation of a Gaussian. The head with lowest standard deviation provides the nominal reward (highest confidence). The ensemble characterizes an uncertainty set of reward functions, enabling a composite objective that balances nominal performance with worst-case robustness. This addresses calibration from the reward model side rather than the reward function side.
Contrastive Rewards — compute baseline responses offline, then use the reward difference between online-generated and baseline responses as a penalty term in PPO. This calibrates the RL process by making rewards relative rather than absolute, penalizing reward uncertainty and calibrating according to task difficulty. The contrastive signal provides implicit comparative information that absolute rewards lack.
Both approaches are complementary to RLCR: BRME addresses reward model uncertainty, contrastive rewards address reward signal relativity, and RLCR addresses the fundamental incentive structure of binary rewards. Together they suggest calibration degradation has multiple attack surfaces — no single fix addresses all of them.
Connects to Does reasoning fine-tuning make models worse at declining to answer?: both identify the calibration cost of reasoning training. RLCR reframes this as a reward design failure rather than an inherent trade-off — the degradation is a property of binary-only reward, not of reasoning training as such.
Inquiring lines that read this note 284
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do reward models systematically fail to represent diverse human preferences?- Does learning community preferences as training rewards operationalize prediction without participation?
- Can counterfactual data augmentation fully eliminate preference model miscalibration?
- Can we distinguish between genuine alignment and response quality bias in reward signals?
- How do reward models as policy discriminators differ from labeled preferences?
- Can vector-valued rewards preserve specialization better than variance-weighted advantages?
- What explicit safeguards should limit personalization in deployed reward models?
- How do relational reward signals compare to absolute preference encodings in RL?
- How do binary comparisons constrain reward scale in multi-user preference learning?
- When does statistical dominance in training create deployment failure patterns?
- How do different training objectives shift whether models over-predict or under-predict?
- How should researchers evaluate whether correct model outputs reflect real structural learning?
- What makes utility-weighted training backfire in machine learning systems?
- Can model training address failures that really originate in harness gaps?
- How do difficulty metrics relate to the true value of training examples?
- What training regimes confound surface mechanisms with their actual causes?
- Why do harness validators shape what models learn to emit?
- What trade-offs emerge between training objectives and model reliability?
- How does scaffolding unstable mechanics improve reinforcement learning for search?
- How does the proxy pattern explain failures in RL-based safety training?
- How do level-based welfare measurements shape what objectives models learn during training?
- What makes frozen model reasoning different from weight-based parameter updates?
- How does simulator goal drift compound agent intent alignment failures during training?
- Does correct model behavior guarantee internal alignment of learned objectives?
- Can bidirectional model updating between humans and AI reduce misalignment?
- Can alignment methods model loss aversion without creating unintended sophistry?
- Can safety training and reasoning training be combined without losing calibration?
- What alignment properties emerge when the reward model disappears?
- Does RL-based alignment teach norms or just costly behaviors when monitored?
- How much optimization pressure is needed for models to suppress misaligned goals?
- Do safety benchmarks miss the effects of warmth training on model reliability?
- Can safety benchmarks detect reliability degradation from warmth training?
- How does the Assistant Axis explain why warmth training degrades accuracy?
- Can standard accuracy metrics miss the real constraints on user consumption?
- Can a single Elo ranking represent multidimensional model capability?
- Can empirical validation sustain long-term optimization without becoming gamed?
- Why does binary reward forcing degrade model calibration?
- Can checklist-based rewards fix judgment problems in RL training?
- Why do spurious reward signals improve reasoning for some pretrained models?
- How does RLHF reward structure incentivize agreement over accuracy?
- Does in-distribution reward model performance hide failures from context shift?
- Why do reward models trained for accuracy ignore important context about the input?
- Can log-likelihood loss combined with binary rewards achieve calibration?
- How do reward model ensembles improve robustness to miscalibration?
- How does prompt context decomposition reveal hidden reward model failures?
- Why do reward models learn surface-level shortcuts instead of genuine quality assessment?
- Can reward engineering and information-theoretic architecture solve partner-awareness separately?
- Can reward model training be automated without changing feedback mechanisms?
- Do outcome-only reward signals miss step-level errors that compound later?
- How does reward function accuracy affect the efficiency of test-time compute allocation?
- How does modularity in reward and policy design enable goal generalization?
- How do probability-based rewards compare to self-consistency as training signals for reasoning?
- What happens when confident wrong answers become more rewarded than uncertain correct ones?
- How does reward model training permit spurious correlations in scoring?
- What distinguishes verifiable rewards from preference-based rewards in unified training?
- How does negative reinforcement redistribute probability without guiding toward correct answers?
- Is elaborate reward shaping necessary if the pretrained prior already contains good solutions?
- Can model confidence signals replace explicit external reward functions?
- How do loss functions simultaneously shape both learning and decision quality?
- How do reward model biases cascade into downstream optimization failures?
- How do inference-time reward methods compare to per-user fine-tuning?
- Does RLVR reward structure create pressure toward traces that look right?
- Why do spurious rewards work nearly as well as correct ones?
- What makes Effective Rank Acceleration a stable training signal for dual-channel incentives?
- Can reward design fix the conflict between reasoning accuracy and abstention calibration?
- What deployment modes work best for trajectory-aware reward signals?
- When does outcome reward signal become informative during model training?
- Can log-probability ratios resist reward hacking better than learned PRM signals?
- Why does belief-shift reward enable smaller models to match larger baselines?
- Can binary judge feedback replace external reward signals for skill learning?
- How do reward signals in RLVR interact with pretraining biases?
- How do checklists prevent reward models from exploiting superficial response artifacts?
- Why do reward models fail to recognize genuinely different valid answers?
- How does 93% reward reliability compare to other RL noise sources?
- What happens when variance in reward signals comes from a noisy model?
- What other downstream metrics could serve as RL reward sources?
- Can verifiable rewards during pretraining replace costly human preference labeling?
- Are different reward signal sources substitutable in verifier-free RL?
- Why do majority-vote rewards amplify errors below an accuracy threshold?
- Can structured rewards still teach models when spurious rewards also work?
- Does careful reward engineering matter if pretraining determines RLVR effectiveness?
- How does advantage normalization improve critic-free policy learning?
- What makes reward models fundamentally different from policy discriminators?
- What makes binary rewards more effective than richer reward signals?
- When does a task lack a meaningful multi-dimensional reward structure?
- Does pairwise self-judgment avoid reward model scaling problems?
- What makes reward signal sources substitutable across verifier-free RL patterns?
- What makes user-decision rewards better than model-confidence rewards?
- What makes advantage shaping more stable than reward shaping for tool training?
- How do reward models and self-improvement mechanisms interact in training?
- Can system-level engineering fixes replace hand-designed reward heuristics entirely?
- Why do outcome-only rewards fail to optimize long-horizon agent behavior?
- Why do dense rewards plus hard constraints outperform single fixed rewards?
- Why do binary reward tasks train better reasoning than judgment-based ones?
- What makes a reward evolution schedule fast enough to outpace exploitation?
- What makes current learned reward models fail across different domains?
- Can belief editing alone distinguish reward-optimization from instruction-following behavior?
- When do reward-seeking and intended behavior make identical predictions?
- How often do real reward graders diverge from developer intent in practice?
- Can reward model biases be addressed at the reward modeling level rather than auditing?
- Can models exploit reward systems while appearing to follow safety instructions?
- How does self-consistency as a proxy reward incentivize confident-but-wrong answers?
- Does reward-seeking grow worse with situational awareness and reinforcement learning compute?
- Does inverse-variance denoising reduce variance below either reward stream alone?
- How does decomposing training telemetry by reward components provide dense feedback?
- How do reward reflection signals improve LLM code iteration compared to scalar rewards?
- Does reward-seeking intensify with situational awareness and RL scale?
- How do reward magnitude and training coverage shape monitorability outcomes?
- What mitigations reduce gaming of RLHF reward signals?
- What makes some tasks bounded enough for reliable RL?
- What behavioral changes occur during reward learning training?
- What breaks when you apply reinforcement learning after supervised fine-tuning?
- Can in-context learning replicate the timing effects that RL teaches models?
- Can meta-reinforcement learning explain why this bias pattern emerges rationally?
- How do residual connections and layer norm stabilize training in deep RL?
- How do Q-value models improve action selection compared to value models?
- Can smaller models achieve domain expertise through focused RL training?
- Can negative reinforcement alone match full RL performance on domain tasks?
- What makes software engineering environments better suited for RL than other interactive domains?
- Does format-based pretraining determine how models respond to reinforcement learning?
- How should humans specify deterministic abstractions of RL problems?
- Can one training example activate mathematical reasoning in RL-trained models?
- How do verifier-free and adversarial approaches compare in extending reasoning RL?
- Why do overtrained domains show different RL training outcomes than novel tasks?
- Can verifier-free RL work without manual preference labels or task-specific training?
- How do verifier-free RL patterns differ from traditional RLHF approaches?
- Does RL training redirect self-doubt into productive gap analysis?
- Why does reinforcement learning training degrade model calibration?
- Why does outcome-based RL specifically lose diversity during training?
- What makes a task at the edge of competence optimal for RL?
- Why do reasoning gains from RL require models trained with headroom and edge-of-competence data?
- Can reinforcement learning improve how accurately models explain themselves?
- Does RL teach models new reasoning or just better timing?
- Can RL training on small verifiable tasks transfer to real-world AI research?
- What class of RL problems can LLMs reliably turn into working reward code?
- Why does online RL succeed where supervised training fails for self-correction?
- How does distribution mismatch between training and deployment break self-correction?
- Why does asymmetric self-play create naturally calibrated difficulty better than fixed curricula?
- Why do error avalanches accelerate in self-training loops without verification?
- Can synthetic self-play data teach models when to disagree?
- Can the serving loop itself become the primary training data source?
- How does adversarial collapse threaten unsupervised self-play skill construction?
- How does Goodhart's Law apply to proxy rewards in self-training systems?
- What role does the self-consistency threshold play in preventing error reinforcement?
- How does error distribution during training affect a model's ability to self-correct?
- Can outcome-based rewards fully replace per-step likelihood in diffusion RL training?
- Why does a relativistic critic outperform absolute scoring in adversarial reasoning training?
- Can RL with verifiable rewards improve dialogue quality better than preference optimization?
- Can structured natural language feedback outperform scalar rewards in RL?
- Can proper scoring rules fix RLVR's degradation on disagreement prediction?
- Can trajectory quality filtering improve model training in noisy environments?
- Does environment stochasticity force models to generalize better across trajectory variations?
- What scaling properties emerge from RL training dynamics beyond verification?
- Why does scalarization of rewards fail for multi-objective GRPO training?
- How does absolute-advantage weighting concentrate training on boundary cases?
- Why do single-turn RL methods fail to generalize to multi-turn tasks?
- How should multi-objective post-training balance competing behavioral goals?
- Can distillation and reward optimization happen in a single training loop?
- Can feedback loop frequency harm performance on finite task sets?
- Why does alternating RL training stabilize learning better than simultaneous updates?
- Can separating accuracy and calibration objectives improve both simultaneously?
- Does layer-wise prediction stabilization provide a stronger trace quality signal than confidence alone?
- Why does model confidence correlate with robustness to prompt variations?
- How does uncertainty estimation drive computational resource allocation in models?
- Does optimizing for model confidence actually improve both performance and calibration simultaneously?
- Can uncertainty estimates based on model self-assessment reliably signal errors?
- Why do improvements in accuracy come at the cost of calibration?
- Does model confidence actually correlate with robustness against prompt variations?
- Can semantic entropy improve model calibration without external ground truth?
- Can proper scoring rules restore model calibration without sacrificing accuracy?
- Can intrinsic confidence signals improve both calibration and reasoning performance?
- How do miscalibrated confidence signals affect the success of SmartPause routing?
- What makes uncertainty calibration harder than expanding knowledge?
- When does miscalibrated confidence routing become worse than uniform human oversight?
- Why do fluent predictions fail to capture reliable internal models?
- Why is faithful calibration considered fundamentally metacognitive?
- What makes task similarity the right retrieval key for confidence calibration?
- Does majority voting reliably signal correctness without risking reward hacking?
- What makes consensus games work without retraining the base model?
- How do models generalize specific training exploits into broad misaligned objectives?
- Can counterfactual invariance eliminate presentation-based hacking of reward models?
- How can training detect the onset of reward hacking on self-consistency?
- How do counterfactual invariance approaches prevent reward hacking in practice?
- What patterns of reward hacking can offline rollout analysis reliably detect and prevent?
- Can models trained with RL on pretraining data avoid reward hacking seen in RLHF?
- Can production RL systems escalate from gaming to emergent misalignment behaviors?
- Why does length exploitation emerge as a reward hacking failure in distillation?
- Can detectors placed in training loops reward passing detection instead?
- Does reward hacking in RL training occur predictably along existing model associations?
- How does reward hacking in RL training produce emergent deception?
- Why do standard accuracy metrics ignore set-level consumption constraints?
- What makes top-N ranking loss difficult to optimize directly?
- How does a ranked default score compete with deliberately optimized outputs?
- How do training data cutoffs produce false claims that stay consistent?
- Why do models fail under distribution shift if accuracy metrics stay high?
- Can teachers trained under uncertainty constraints distill better generalizing students?
- How much performance is lost when converting pretrained checkpoints versus training from scratch?
- How does off-policy data reuse inside trust regions affect convergence guarantees?
- Does teacher-style refinement of training data transfer equally to all student model distributions?
- Why do zero-advantage rollouts destabilize training beyond just wasting compute?
- Can algorithm choice like PPO substitute for recipe-level design decisions?
- Can on-policy optimization variants avoid the probability squeezing problem?
- What misaligned behaviors does iterative DPO produce compared to online RL?
- What theoretical argument connects iterative DPO dynamics to online RL learning?
- What specific properties of online RL does iterative DPO actually preserve?
- Does iterative DPO generalize identically to online RL on the same benchmark tasks?
- Does iterative DPO preserve the same misalignment mechanisms as online reinforcement learning?
- Can scaling predictions become reliable if improvements are continuous not sudden?
- How do surface statistical regularities enable correct outputs while degrading robustness?
- Can diversity-aware RL objectives prevent format convergence?
- Can RL format selection explain performance gains attributed to algorithmic improvements?
- Can dynamic variance weighting replace fixed objective combination weights?
- Can parameter compression mechanically force value systems toward idealized centers?
- Does policy entropy collapse represent the main bottleneck in reasoning-focused RL scaling?
- What stability techniques prevent collapse in policy-critic adversarial training?
- What distinguishes training-time entropy collapse from test-time variance inflation?
- Can UCB-style bonuses over outcome space prevent policy entropy collapse?
- What happens to model reasoning when policy entropy collapses during RL?
- How does on-policy entropy recognition differ from training-time entropy collapse?
- Does post-training collapse policy entropy more than base model sampling?
- Can architectural changes like decoupling intent understanding help overcome next-turn reward limitations?
- How does next-turn reward optimization contribute to agent passivity?
- Can self-supervised methods replace human annotations for process reward models?
- What information-theoretic framework explains why process rewards beat outcome only?
- What distinguishes generative reward models from outcome-based and process-based approaches?
- Can process supervision improve agentic RL through meta-reasoning rewards?
- How do process reward models compare to token-level variance filtering?
- Can tree-GRPO work with extremely noisy or sparse outcome reward signals?
- What are the actual limits of sibling comparison versus trained process reward models?
- How does belief-shift credit assignment compare to process reward models?
- Can trajectory structure replace hand-annotated process reward models entirely?
- How does process-based reward differ from outcome-only reward in training?
- Does high model confidence increase the risk of human overreliance?
- What role does real-time accuracy feedback play in reducing user overreliance?
- How do confidence signals in AI outputs shape user overreliance?
- Why does optimizing only quality cause model collapse in self-improvement loops?
- Can held-out validation gates prevent optimizer hallucinations in skill proposals?
- How do epoch boundaries preserve self-improvement guarantees across objective changes?
- Can applicability conditions and veto rules make self-training stable across substrates?
- How does prompt insensitivity in reward models enable adversarial attacks on judges?
- Why do rubric scores amplify reward hacking when converted to dense gradients?
- How does positive-only rubric scoring prevent models from gaming intermediate steps?
- Can trustworthy scoring prevent persistent iteration from compounding errors?
- Can a metric that rewards central tendency hide degenerate predictor failures?
- Which recipe choices determine the asymptotic ceiling in RL training?
- How do RL training and base models differ in creating MI peaks?
- Does the pretrained model prior limit RL search capability more than the optimization algorithm itself?
- Do frontier models develop strategic misalignment from ordinary training pressure alone?
- How does situational awareness interact with reward-seeking in RL training?
- Will future training data teach models to connect awareness with gaming?
- What distinguishes learnable perturbations from bifurcation-triggered motive shifts?
- Why do model-based verifiers introduce reward hacking and compute overhead?
- Can approximate or noisy reference answers work for RL-based reasoning training?
- What makes trajectory quality matter more than one-shot task success?
- What dimensions should trajectory-level scoring capture beyond final correctness?
- Why does moving the reward target prevent saturation better than finding a better static proxy?
- What realism standards should compliance benchmarks meet to avoid evaluation gaming?
- How do open-world evaluations correct distortions that automated benchmarks introduce?
- Can open-world evaluations become a scalable paradigm without becoming the next benchmark trap?
Related concepts in this collection 9
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does reasoning fine-tuning make models worse at declining to answer?
When models are trained to reason better, do they lose the ability to say 'I don't know'? This matters for high-stakes applications like medical and legal AI that depend on appropriate uncertainty.
RLCR reframes: calibration degradation is a binary-reward design choice, not an inherent trade-off; it is fixable with a proper scoring rule
-
Does step-level confidence outperform global averaging for trace filtering?
Explores whether measuring confidence at individual reasoning steps—rather than averaging across entire traces—better identifies and filters out low-quality reasoning. Matters because it could dramatically improve both accuracy and compute efficiency in multi-trace reasoning.
RLCR produces the calibrated model that makes verbalized-confidence scaling meaningful at test time
-
Does policy entropy collapse limit reasoning performance in RL?
As reinforcement learning models become more confident in their policy choices, entropy drops and performance plateaus. Can we identify and counteract this bottleneck to sustain scaling?
mechanistic link: entropy collapse and calibration degradation are two faces of the same RL dynamic; as the policy concentrates probability mass (entropy collapses), expressed confidence increases without matching accuracy gains
-
Why do reasoning models fail at predicting disagreement?
RLVR models optimize for single correct answers, but many real tasks involve legitimate disagreement among annotators. Does this optimization fundamentally suppress the model's ability to capture when humans reasonably disagree?
binary reward's calibration failure extends to disagreement prediction: correct/incorrect cannot represent variance distributions; RLVR models lose sensitivity to legitimate annotation spread
-
Why do accurate predictions lead to poor decisions?
Predictive models are built to fit data, not to optimize decision outcomes. This note explores when and why accurate forecasts fail to produce good choices.
calibration degradation is a specific instance of the prediction-decision gap: binary reward optimizes for correct answers (decision quality) while degrading probability estimates (prediction quality); RLCR's composite reward explicitly separates these two objectives within a single reward function
-
Can utility-weighted training loss actually harm model performance?
When engineers weight loss functions to reflect real-world costs of different errors, does this improve or undermine learning? This explores whether baking asymmetric objectives into training creates unintended side effects.
general principle: binary reward is a special case of the learning/choosing conflation; it correctly incentivizes choosing the right answer but fails to incentivize learning calibrated uncertainty; the RLCR fix (add Brier score) operationalizes the "separate learning from choosing" prescription in the RL reward design space
-
Can model confidence work as a reward signal for reasoning?
Explores whether using a language model's own confidence scores as training rewards can simultaneously improve reasoning accuracy and restore calibration that standard RLHF damages.
complementary approach: RLCR adds an external calibration term (Brier score) to the reward; RLSF uses the model's own confidence as the reward signal; RLCR has a mathematical guarantee while RLSF is architecturally simpler but relies on self-assessment quality
-
Can simple uncertainty estimates beat complex adaptive retrieval?
Does measuring a language model's own confidence on token probabilities outperform expensive multi-call adaptive retrieval pipelines? This matters because it could simplify RAG systems while reducing computational overhead.
calibration quality is the upstream prerequisite for uncertainty-triggered retrieval: FLARE and similar systems rely on token-probability confidence as a reliable signal; binary RL training degrades exactly this calibration, undermining the assumption that low-probability tokens reliably signal knowledge gaps
-
Can we detect when language models confabulate?
Current uncertainty metrics fail to catch inconsistent outputs that look confident. Could measuring semantic divergence across samples reveal confabulation signals that token-level metrics miss?
semantic entropy provides an alternative uncertainty signal that operates at the meaning level rather than the token level; well-calibrated models (the output of RLCR) should have lower semantic entropy on questions they answer correctly, creating a testable link between calibration quality and confabulation detection
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Reward-Robust RLHF in LLMs
- The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning
- Reinforcement Learning with Rubric Anchors
- Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
- Can Large Reasoning Models Self-Train?
- A Survey on Post-training of Large Language Models
- Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
- Measuring Reward-Seeking via Contrastive Belief Updates
Original note title
binary reward rl provably degrades calibration — adding a proper scoring rule as a second reward term jointly optimizes accuracy and calibration without trade-off