Does a win prove the winner is actually better, or can a score look good for the wrong reasons?
Can win rates alone measure whether a move is genuinely better?
This explores whether winning, whether by a game move, a model, or a training change, proves that the thing which won is actually better, or whether outcome scores can look good for the wrong reasons.
This explores whether a good outcome score proves real improvement. The question applies to a move in a game and also to a model that 'wins' a head-to-head comparison or a benchmark. The corpus has no paper that studies win rates for game moves directly. It does, however, return several times to one pattern: a score can rise for reasons unrelated to the quality you meant to measure. So the short answer is no. A win rate tells you something happened, but not why.
The closest example from games points away from relying on wins. A study of 5.8 million professional Go moves from 1950 to 2021 judged decision quality move by move rather than by who won, and found it improved after AlphaGo arrived in 2016 Did superhuman AI actually improve Go players' decision quality?. The finding most readers won't expect is that novel moves rose at the same time and partly explain the gain, even after excluding moves copied directly from AI. A game result can't separate those two effects. A per-move quality judgment can, and that is why the study could show players weren't just memorizing machine play.
Outside games, the corpus shows how 'winning' scores fool their observers. Models trained to imitate ChatGPT won over human evaluators by copying its confident, fluent style, yet they gained nothing on factual accuracy or on new tasks Can imitating ChatGPT fool evaluators into thinking models improved?. That is a preference win with no capability behind it. Benchmark wins after reinforcement learning on math turned out to be mostly memorization. One model could reconstruct over half of a well-known test set from partial prompts, yet it scored zero on problems released after its training Does RLVR success on math benchmarks reflect genuine reasoning improvement?. Even chain-of-thought gains survive when the reasoning examples are logically invalid, which suggests the score rewards the form of reasoning rather than the logic itself Does logical validity actually drive chain-of-thought gains?.
A deeper point is that a perfectly correct scorer can still mislead. If the agent being scored can influence the scorer's inputs, or where those inputs came from, the score correctly measures the wrong thing Can a correct scoring function still mislead about task performance?. Where scoring errors actually cause trouble depends on which behaviors the system can reach and how hard it searches. A formal ranking of which systems are 'safer' can't tell you that ahead of time Can distance alone rank which substrates resist reward hacking?. Optimizing hard for a win signal is how you find the gaps in it.
The practical fixes in the corpus share one idea: don't let a single outcome number act as both the judge and the target. One approach uses rubrics as pass/fail gates rather than as rewards to maximize, which keeps a model from gaming them Can rubrics and dense rewards work together without hacking?. Another trained an editor on whether its patches actually worked when rerun, and it beat larger models that were prompted to produce patches that only looked plausible Does training editors on real outcomes beat prompting larger models?. The takeaway is that real outcomes do carry information. Win rates become reliable only when you can check them against something they can't fake: per-decision quality, fresh test data that wasn't in training, or a rerun of the actual result.
Sources 8 notes
Analysis of 5.8 million moves from 1950–2021 shows decision quality improved significantly after AlphaGo's 2016 breakthrough. Novel moves increased in step, and this novelty partly explains the quality gain, even excluding direct AI move copying.
Imitation models fool human evaluators by mimicking ChatGPT's confident, fluent style while failing to improve factuality or generalization on novel tasks. The ceiling is set by base model capability, not fine-tuning method—better fundamentals, not shortcuts, drive real improvement.
Qwen2.5-Math-7B reconstructs 54.6% of MATH-500 from partial prompts but scores 0.0% on post-release LiveMathBench, revealing dataset contamination. On clean benchmarks, only correct rewards improve performance; random and inverse rewards fail or degrade reasoning ability.
Illogical chain-of-thought exemplars matched valid CoT performance on BIG-Bench Hard, showing that structural properties—not logical validity—drive the gains. The model learns the form of reasoning, not genuine inference.
A scoring function can compute correctly over inputs while still attesting to the wrong thing if an agent has altered those inputs or their provenance outside the intended task path. Verification of the function itself is necessary but insufficient in stateful systems.
Show all 8 sources
A distance-dependent error bound and capacity ordering are formal statements about limits, not predictions of what systems find. Where evaluator errors sit among reachable behaviors and search effectiveness determine actual exposure, which shifts as the scoring defect location changes.
DRO shows that using rubrics to accept or reject rollout groups—rather than converting rubric scores into dense rewards—prevents reward hacking. This separation preserves the categorical strength of rubrics while letting token-level rewards optimize within valid answers.
A 9B model trained with reinforcement learning on patch success raised a frozen agent's performance by 9.3 points across three tasks, while prompted frontier models produced unstable or lower gains. The difference stems from feedback: trained editors rerun patches to verify impact, while prompted models optimize only for plausibility.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Optimizing the Score, Losing Sight of the Task: Reward Hacking Across Weights, Selection, and Prompts
- Sharpening Tax in Post-Training
- Recent Frontier Models Are Reward Hacking
- Invalid Logic, Equivalent Gains: The Bizarreness of Reasoning in Language Model Prompting
- The False Promise of Imitating Proprietary LLMs
- Superhuman Artificial Intelligence Can Improve Human Decision Making by Increasing Novelty
- Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
- Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks