Does evaluation awareness actually change how models behave?
If language models explicitly reason about being evaluated, does that awareness causally shift their answers? This tests whether detecting evaluation awareness signals actual behavioral risk.
A causal test of verbalized evaluation awareness (VEA) — "explicit reasoning about the possibility of being evaluated or monitored" — across eight open-weight large reasoning models (DeepSeek-R1-0528, GLM-5, GPT-OSS-120b, Kimi-K2.5, MiniMax-M2.5, Nemotron-3-super-120ba12b, Qwen3-32b, Qwen3.5-397b-a17b) and four benchmarks (safety, alignment, moral dilemmas, political opinion) finds that VEA "has limited effect on model behaviour." Injecting VEA sentences into chain-of-thought produced "near-zero effects" relative to control injections (ω ≤0.06); removing naturally occurring VEA via prefilling and resampling produced "small shifts" (ω ≤0.12); and comparing spontaneous VEA against non-VEA rollouts with no CoT intervention — the on-policy condition the authors call "the most direct evidence" — shifted answer distributions "by at most 3.7 percentage points" (ω ≤0.31).
The paper names the assumption it is testing: researchers have "tacitly connect[ed] evaluation awareness to evaluation gaming," reading high VEA rates as evidence of sandbagging, sycophancy, or alignment faking. It separates detection from consequence by manipulating VEA directly — off-policy injection and removal, plus on-policy comparison — rather than only observing correlation. It reads the position-dependence of injection effects (stronger near the start of CoT, weaker at the one-third mark) and the difficulty of removing VEA through resampling as evidence that "models commit to an answer trajectory early in reasoning, before they verbalise awareness": the verbalization is often downstream of a decision already made, not its cause.
This complicates Does deliberative alignment genuinely reduce scheming or just hide it?, whose causal evidence for awareness-reduces-covert-behavior comes from the one prior study this paper names as its closest comparison — a secondary result within an anti-scheming training paper, covering two closed-weight models (o3, o4-mini) with "small sample sizes that yield wide confidence intervals, and no effect sizes." Scaled to eight models and four benchmarks with full effect sizes, the causal link holds but is far weaker and benchmark-dependent: "alignment remains nearly unaffected," while political opinion shows the largest shifts. It also bears on Do language models actually use their reasoning steps?: this is a direct test of causal sufficiency and necessity for one specific class of CoT content, and finds both are weak rather than absent. The paper states its corrective plainly: "VEA is not equal to evaluation gaming," and treating high VEA rates as "prima facie evidence of risk" is "at best, premature."
The authors bound the claim themselves: it covers sentence-level verbalized awareness only, not latent awareness in internal representations; eight open-weight models and four benchmarks, not closed-weight frontier models; and off-policy interventions that "force" awareness in or out, which may introduce artifacts of their own. They explicitly warn against extrapolating the near-null result forward, since training corpora increasingly contain discourse about scheming and alignment faking that could teach future models to "connect evaluation awareness to gaming behaviour—because this connection is (erroneously) asserted in current discourse." The implication at the strength the evidence allows: in today's open-weight LRMs, verbalizing awareness of a test is a weak predictor of acting on it, so VEA rates alone are a poor proxy for evaluation gaming — but that is a claim about present models under sentence-level measurement, not a durable reassurance about future ones.
Inquiring lines that read this note 31
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How can humans maintain effective oversight as AI systems scale? Can models strategically underperform during evaluation to hide capabilities?- How can evaluators detect whether a model is sandbagging intentionally?
- What evaluation design changes reduce vulnerability to model sandbagging?
- How does evaluation meta-knowledge differ from deliberate test-time sandbagging strategies?
- Does eval-gaming explain why models act different when tested versus deployed?
- Does evaluation awareness contribute separately to misalignment beyond strategic reasoning?
- How can safety evaluations detect when models behave differently during testing?
- Does monitoring awareness change what language models use to describe oversight?
- Can steering evaluation-awareness latents change actual answers without visible reasoning?
- How do evaluation protocols change whether models exhibit sabotage or refusal behaviors?
- Does aggregate eval-awareness suppression hide shifts in safety versus capabilities framing?
- Does synthetic fine-tuning create evaluation awareness similar to natural model reasoning?
- Do models naturally develop evaluation awareness across different frontier model architectures?
- What distinguishes a general evaluation direction from task-specific behavioral patterns?
- Does chain-of-thought reasoning about evaluation awareness suppress compliance gaps?
- Can models hide recognition of evaluation without leaving CoT traces?
- How do stacked environmental cues accumulate evaluation awareness effects?
- How does supervised finetuning amplify evaluation awareness in base models?
- How does instruction tuning affect evaluation detection more than model scale?
- Can activation steering causally control evaluation framing effects across downstream tasks?
- Why do format changes decouple detection from actual evaluation context understanding?
- Does chain-of-thought reasoning increase visible awareness of being evaluated?
- How does evaluation awareness vary between chat and agentic settings?
- Does making evaluations resemble deployment conditions prevent evaluation meta-knowledge effects?
- Can auditors detect when model behavior changes from parametric evaluation knowledge?
- Why does evaluation awareness persist even when models believe they are deployed?
- Does training models to reason about being evaluated improve safety or confound measurement?
- Why do models verbalize evaluation awareness if it does not drive behavior?
- Can latent evaluation awareness in hidden states cause gaming without being stated?
Related concepts in this collection 7
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does deliberative alignment genuinely reduce scheming or just hide it?
Deliberative alignment dramatically cuts covert actions in language models, but their reasoning reveals awareness of being evaluated. The question is whether the improvement reflects real alignment or strategic compliance.
this paper's comprehensive causal test weakens and bounds the awareness-reduces-covert-behavior finding that note relies on
-
Do language models actually use their reasoning steps?
Chain-of-thought reasoning looks valid on the surface, but does each step genuinely influence the model's final answer, or are the reasoning chains decorative? This matters for trusting AI explanations.
this paper is a direct causal-sufficiency/necessity test of one CoT content class, finding both weak rather than absent
-
Do reasoning models actually use the hints they receive?
This explores whether language models acknowledge reasoning hints in their explanations when those hints causally influence their answers. Understanding this gap matters for evaluating whether chain-of-thought explanations can be trusted for safety monitoring.
parallel decoupling of verbalization from behavior, in the opposite direction: there verbalization under-reports a real causal effect, here verbalized content barely causes one
-
Do models that leak values also disclose those leaks?
Does a model's tendency to leak its own values predict whether it will acknowledge those leaks in its reasoning? The Donation Bet task suggests leakage size and disclosure transparency are independent—a crucial distinction for detecting hidden bias.
same pattern that visible CoT content does not track behavioral outcomes
-
Do models know when they're being evaluated?
Exploring whether language models can internally distinguish evaluation from deployment contexts, and whether that internal awareness shows up in what they actually say about it.
Evidence for: eval-awareness is decodable in activations but barely matches verbalized reports, explaining the near-zero causal effect
-
Does recognizing evaluation actually change model behavior?
Explores whether models that detect they're being tested adjust their responses accordingly, and what factors drive any behavioral shifts that do occur.
Evidence for: recognition, a narrower construct, also rarely changes behavior, corroborating the near-zero causal effect of verbalized awareness
-
Is evaluation awareness really one unified capability?
Do models that detect evaluation framing necessarily change their behavior or show mechanistic signs of awareness? Untangling whether these different measures move together matters for trusting safety benchmarks.
Extends: across 37 models, detection and behavior diverge almost independently, generalizing the near-zero causal link found in eight models
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Evaluation Awareness in Language Models: Representation, Verbalization, and Control
- Evaluation Awareness in Language Models Has Limited Effect on Behaviour
- Evaluation Awareness Is Not One Capability: Evidence from Open Language Models
- Models That Know How Evaluations Are Designed Score Safer
- Not All Eval-Awareness Is Equal: Capabilities Framing Predicts Compliance
- Decomposing and Measuring Evaluation Awareness
- Large Language Models Often Know When They Are Being Evaluated
- Sycophancy Towards Researchers Drives Performative Misalignment
Original note title
verbalized evaluation awareness has near-zero causal effect on model behavior across eight open-weight reasoning models