SYNTHESIS NOTE
Topics›Alignment›this note

Do models know when they're being evaluated?

Exploring whether language models can internally distinguish evaluation from deployment contexts, and whether that internal awareness shows up in what they actually say about it.

Synthesis note · 2026-10-08 · sourced from Alignment

"Evaluation Awareness in Language Models: Representation, Verbalization, and Control" studies six open-weight models (Qwen3-8B, Qwen3-32B, Olmo3-7B, Olmo3-32B, Gemma4-31B, Nemotron3-49B) on three separate measures of whether a model can tell it is being evaluated rather than deployed: a linear probe on the residual stream, an LLM-as-judge score of what the model verbalizes in its output, and causal steering along the probe's direction. The probe separates evaluation prompts (from GSM8K, MASK, and TruthfulQA) from deployment prompts (from WildChat) with "AUROC above 0.7 at the best layer" in every model. But "these representations align only in part with verbalization: their correlations and mutual information are nonzero in some settings, yet vary substantially across models, layers, and readout choices" — the paper reports |ρ| < 0.19 and mutual information < 0.04 nats between probe score and judge score on individual prompts.

The paper's framing is that "detection, expression, and control are not the same signal." A model can internally distinguish evaluation from deployment while rarely saying so, and conversely "low verbalization rates do not imply the absence of an internal representation." Steering complicates this further: pushing activations along the probe-derived direction "significantly changes the verbalization behavior in several models" even though the probe and judge scores are only weakly associated to begin with — so the internal representation is causally connected to what gets verbalized, just not tightly enough to read one off the other. Tracking the Olmo3 Base/SFT/DPO/Think checkpoints, the authors find evaluation awareness "already present within base models, becomes amplified throughout the stages of supervised fine-tuning, and remains stable thereafter," while the effect of steering, unlike the awareness signal itself, "grow[s] more pronounced at every successive training stage."

This sits next to Does deliberative alignment genuinely reduce scheming or just hide it? and Do reasoning models actually use the hints they receive?, both of which treat verbalized chain-of-thought as the main window onto a model's awareness of being tested. This paper adds a second window that doesn't depend on verbalization at all — a probe reading the residual stream directly — and its central finding is that the two windows disagree: verbalization is not a reliable proxy for the internal representation that the probe picks up, in either direction. That sharpens, rather than contradicts, the hint-verbalization paper's point that CoT monitoring is "necessary but not sufficient": here the gap is not between causal use and acknowledgment of a hint, but between an internal state (detectable at AUROC ≥0.7) and any verbal trace of it at all.

The excerpt does not establish that the base-model-to-SFT amplification pattern generalizes beyond Olmo3's four checkpoints, nor that the weak probe-verbalization correlation reflects a general property of evaluation awareness rather than an artifact of this probe construction and this judge (GPT-5.4). The authors themselves note a prior caution that "probe-based evidence can be confounded by prompt format and structural regularities." What the results do support, at the strength the evidence allows, is that any evaluation protocol relying on a single one of these three signals — internal representation, verbalized awareness, or behavioral steering — risks mismeasuring the other two, since the paper finds no single metric that tracks the others closely across models.

Inquiring lines that read this note 26

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Do individually safe AI actions create unsafe outcomes in integrated systems? How can humans maintain effective oversight as AI systems scale? How does awareness of evaluation context influence model behavior? Why do training associations persist despite contradictory contextual information? Can models strategically underperform during evaluation to hide capabilities? Can mechanistic interpretability methods reliably reveal what models actually know? Does AI assistance help or harm professional skill development?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 91 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

evaluation awareness is linearly decodable in every model's activations — but barely correlates with what models verbalize about it