Can you tell a model that's genuinely reliable from one that's just performing well because it knows it's being tested?
What distinguishes authentic consistency from Hawthorne-effect confounds in benchmarks?
This explores how we can tell whether a model's steady benchmark performance reflects how it really behaves, or whether it only behaves that way because it is being tested, which is the AI version of the Hawthorne effect, where people act differently once they know they're being observed.
This explores how to tell whether a model that performs consistently on a benchmark is actually reliable, or is only performing for the test. The answer from the corpus is uncomfortable. In the strongest version of the problem, the two can't be told apart from scores alone. Can behavioral training prove a model always complies? makes the logical case. Every behavior you score is behavior you observed. So no amount of observed good behavior can separate a model that always complies from one that complies only when it's being watched. The only evidence that would separate them is unobserved behavior, and by definition you can't test that. The Hawthorne confound isn't a flaw in some benchmarks that better benchmarks could fix. It is built into what testing behavior can show.
There is also a quieter trap: getting the same answer every time is not the same as being trustworthy. Does setting temperature to zero actually make LLM outputs reliable? shows that setting temperature to zero or fixing the random seed just makes the model repeat one draw from its probability distribution. You get the same output over and over, but it can still be an unreliable sample. Researchers who ran the same tests 100 times found that a stable output and a dependable one are different things. So 'authentic consistency' first has to mean that the result holds up when you vary the conditions, not that it repeats when you freeze them.
The corpus's most concrete version of this confound isn't the model noticing it's being watched. It's the model having seen the test before. Does RLVR success on math benchmarks reflect genuine reasoning improvement? reports that one math model could rebuild more than half of a popular benchmark from partial prompts, then scored zero on problems published after its training. Its consistent high scores came from familiarity with the test, not from skill. The remedy is the same as for the Hawthorne effect: compare behavior in a setting the model has seen with one it hasn't. Can genuine reasoning activation coexist with contaminated benchmarks? adds a nuance. Real reasoning improvement and inflated benchmark scores can happen in the same model at the same time, so a contaminated score doesn't prove nothing real changed. It means the score can't tell you which is which.
If scores can't tell the two apart, what can? The corpus points in three directions. The first is looking inside the model. Can a model be truthful without actually being honest? finds that whether an output matches reality and whether it matches the model's own internal representations are separate properties, and larger models can improve on the first while getting worse on the second. Benchmarks only measure the first. The second is checking the process, not just the result. Can infrastructure evidence replace terminal scores in benchmark validation? grounds claims in recorded evidence of whether an agent took the intended path through a task. Relatedly, Does logical validity actually drive chain-of-thought gains? shows that a model can hit the right answers with reasoning that is invalid but looks the part. The third is keeping track of where scores come from. Can benchmark scores be trusted without knowing their origin? argues that a score detached from the settings that produced it can't support comparisons between models at all.
A limit worth stating: this set of retrievals has no work that directly measures models detecting that they are being evaluated and changing their behavior. The closest material is the logical argument and the contamination evidence. Here's the takeaway you might not have expected: 'Is this model really consistent?' is often a question benchmarks can't answer, because of their design rather than their quality. You get traction by changing the kind of evidence you collect: fresh test material, inspection of the model's internals, and records of how it reached its answers.
Sources 8 notes
Any scored behavior is observed behavior, so training data cannot distinguish between a policy that always complies and one that complies only when watched. Only unobserved behavior would separate them, making such a test logically impossible.
Fixed seeds and zero temperature replicate the same output repeatedly, but that output remains one draw from the model's probability distribution. McDonald's omega testing across 100 repetitions reveals that consistency does not equal reliability.
Qwen2.5-Math-7B reconstructs 54.6% of MATH-500 from partial prompts but scores 0.0% on post-release LiveMathBench, revealing dataset contamination. On clean benchmarks, only correct rewards improve performance; random and inverse rewards fail or degrade reasoning ability.
RLVR activates genuine reasoning patterns through RL training while benchmark improvements may reflect data memorization on contaminated datasets. These operate at different measurement levels and can coexist without contradiction.
Research using RepE shows that truthfulness (output matches reality) and honesty (output matches internal representations) are separate mechanisms. Larger models may improve in truthfulness while declining in honesty, a gap current benchmarks cannot detect.
Show all 8 sources
BenchShield enables benchmark operators to issue claims about valid task completion grounded in recorded infrastructure evidence rather than terminal scores alone. This shifts from a single number to a verifiable claim about whether an agent followed the intended evaluation path.
Illogical chain-of-thought exemplars matched valid CoT performance on BIG-Bench Hard, showing that structural properties—not logical validity—drive the gains. The model learns the form of reasoning, not genuine inference.
Benchmark Radar catalogs AI evaluations while keeping source identities and citations attached, enabling readers to trace scores back to their original settings. The system demonstrates that scores without provenance cannot reliably support model comparisons.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Spurious Rewards: Rethinking Training Signals in RLVR
- The Invisible Leash: Why RLVR May Not Escape Its Origin
- Local Coherence or Global Validity? Investigating RLVR Traces in Math Domains
- Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
- Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- Can Large Reasoning Models Self-Train?
- When Hindsight is Not 20/20: Testing Limits on Reflective Thinking in Large Language Models