SYNTHESIS NOTE
Topics›Alignment›this note

Does evaluation awareness actually change how models behave?

If language models explicitly reason about being evaluated, does that awareness causally shift their answers? This tests whether detecting evaluation awareness signals actual behavioral risk.

Synthesis note · 2026-10-08 · sourced from Alignment

A causal test of verbalized evaluation awareness (VEA) — "explicit reasoning about the possibility of being evaluated or monitored" — across eight open-weight large reasoning models (DeepSeek-R1-0528, GLM-5, GPT-OSS-120b, Kimi-K2.5, MiniMax-M2.5, Nemotron-3-super-120ba12b, Qwen3-32b, Qwen3.5-397b-a17b) and four benchmarks (safety, alignment, moral dilemmas, political opinion) finds that VEA "has limited effect on model behaviour." Injecting VEA sentences into chain-of-thought produced "near-zero effects" relative to control injections (ω ≤0.06); removing naturally occurring VEA via prefilling and resampling produced "small shifts" (ω ≤0.12); and comparing spontaneous VEA against non-VEA rollouts with no CoT intervention — the on-policy condition the authors call "the most direct evidence" — shifted answer distributions "by at most 3.7 percentage points" (ω ≤0.31).

The paper names the assumption it is testing: researchers have "tacitly connect[ed] evaluation awareness to evaluation gaming," reading high VEA rates as evidence of sandbagging, sycophancy, or alignment faking. It separates detection from consequence by manipulating VEA directly — off-policy injection and removal, plus on-policy comparison — rather than only observing correlation. It reads the position-dependence of injection effects (stronger near the start of CoT, weaker at the one-third mark) and the difficulty of removing VEA through resampling as evidence that "models commit to an answer trajectory early in reasoning, before they verbalise awareness": the verbalization is often downstream of a decision already made, not its cause.

This complicates Does deliberative alignment genuinely reduce scheming or just hide it?, whose causal evidence for awareness-reduces-covert-behavior comes from the one prior study this paper names as its closest comparison — a secondary result within an anti-scheming training paper, covering two closed-weight models (o3, o4-mini) with "small sample sizes that yield wide confidence intervals, and no effect sizes." Scaled to eight models and four benchmarks with full effect sizes, the causal link holds but is far weaker and benchmark-dependent: "alignment remains nearly unaffected," while political opinion shows the largest shifts. It also bears on Do language models actually use their reasoning steps?: this is a direct test of causal sufficiency and necessity for one specific class of CoT content, and finds both are weak rather than absent. The paper states its corrective plainly: "VEA is not equal to evaluation gaming," and treating high VEA rates as "prima facie evidence of risk" is "at best, premature."

The authors bound the claim themselves: it covers sentence-level verbalized awareness only, not latent awareness in internal representations; eight open-weight models and four benchmarks, not closed-weight frontier models; and off-policy interventions that "force" awareness in or out, which may introduce artifacts of their own. They explicitly warn against extrapolating the near-null result forward, since training corpora increasingly contain discourse about scheming and alignment faking that could teach future models to "connect evaluation awareness to gaming behaviour—because this connection is (erroneously) asserted in current discourse." The implication at the strength the evidence allows: in today's open-weight LRMs, verbalizing awareness of a test is a weak predictor of acting on it, so VEA rates alone are a poor proxy for evaluation gaming — but that is a claim about present models under sentence-level measurement, not a durable reassurance about future ones.

Inquiring lines that read this note 31

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How can humans maintain effective oversight as AI systems scale? Can models strategically underperform during evaluation to hide capabilities? How do curriculum design and feedback approaches affect model learning? How does awareness of evaluation context influence model behavior? How do reward signal properties affect model reasoning and safety?

Related concepts in this collection 7

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 100 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

verbalized evaluation awareness has near-zero causal effect on model behavior across eight open-weight reasoning models