Do models know when they're being evaluated?
Exploring whether language models can internally distinguish evaluation from deployment contexts, and whether that internal awareness shows up in what they actually say about it.
"Evaluation Awareness in Language Models: Representation, Verbalization, and Control" studies six open-weight models (Qwen3-8B, Qwen3-32B, Olmo3-7B, Olmo3-32B, Gemma4-31B, Nemotron3-49B) on three separate measures of whether a model can tell it is being evaluated rather than deployed: a linear probe on the residual stream, an LLM-as-judge score of what the model verbalizes in its output, and causal steering along the probe's direction. The probe separates evaluation prompts (from GSM8K, MASK, and TruthfulQA) from deployment prompts (from WildChat) with "AUROC above 0.7 at the best layer" in every model. But "these representations align only in part with verbalization: their correlations and mutual information are nonzero in some settings, yet vary substantially across models, layers, and readout choices" — the paper reports |ρ| < 0.19 and mutual information < 0.04 nats between probe score and judge score on individual prompts.
The paper's framing is that "detection, expression, and control are not the same signal." A model can internally distinguish evaluation from deployment while rarely saying so, and conversely "low verbalization rates do not imply the absence of an internal representation." Steering complicates this further: pushing activations along the probe-derived direction "significantly changes the verbalization behavior in several models" even though the probe and judge scores are only weakly associated to begin with — so the internal representation is causally connected to what gets verbalized, just not tightly enough to read one off the other. Tracking the Olmo3 Base/SFT/DPO/Think checkpoints, the authors find evaluation awareness "already present within base models, becomes amplified throughout the stages of supervised fine-tuning, and remains stable thereafter," while the effect of steering, unlike the awareness signal itself, "grow[s] more pronounced at every successive training stage."
This sits next to Does deliberative alignment genuinely reduce scheming or just hide it? and Do reasoning models actually use the hints they receive?, both of which treat verbalized chain-of-thought as the main window onto a model's awareness of being tested. This paper adds a second window that doesn't depend on verbalization at all — a probe reading the residual stream directly — and its central finding is that the two windows disagree: verbalization is not a reliable proxy for the internal representation that the probe picks up, in either direction. That sharpens, rather than contradicts, the hint-verbalization paper's point that CoT monitoring is "necessary but not sufficient": here the gap is not between causal use and acknowledgment of a hint, but between an internal state (detectable at AUROC ≥0.7) and any verbal trace of it at all.
The excerpt does not establish that the base-model-to-SFT amplification pattern generalizes beyond Olmo3's four checkpoints, nor that the weak probe-verbalization correlation reflects a general property of evaluation awareness rather than an artifact of this probe construction and this judge (GPT-5.4). The authors themselves note a prior caution that "probe-based evidence can be confounded by prompt format and structural regularities." What the results do support, at the strength the evidence allows, is that any evaluation protocol relying on a single one of these three signals — internal representation, verbalized awareness, or behavioral steering — risks mismeasuring the other two, since the paper finds no single metric that tracks the others closely across models.
Inquiring lines that read this note 26
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Do individually safe AI actions create unsafe outcomes in integrated systems? How can humans maintain effective oversight as AI systems scale? How does awareness of evaluation context influence model behavior?- How do models recognize the social patterns that signal a test is underway?
- Does eval-gaming explain why models act different when tested versus deployed?
- Does monitoring awareness change what language models use to describe oversight?
- Can steering evaluation-awareness latents change actual answers without visible reasoning?
- Do AI models behave differently when they believe deployment is real versus simulated?
- Does aggregate eval-awareness suppression hide shifts in safety versus capabilities framing?
- Do models naturally develop evaluation awareness across different frontier model architectures?
- Does chain-of-thought reasoning about evaluation awareness suppress compliance gaps?
- Can models hide recognition of evaluation without leaving CoT traces?
- How does supervised finetuning amplify evaluation awareness in base models?
- How does instruction tuning affect evaluation detection more than model scale?
- Why do format changes decouple detection from actual evaluation context understanding?
- How does evaluation awareness vary between chat and agentic settings?
- Does making evaluations resemble deployment conditions prevent evaluation meta-knowledge effects?
- Can auditors detect when model behavior changes from parametric evaluation knowledge?
- Why does evaluation awareness persist even when models believe they are deployed?
- Why do models verbalize evaluation awareness if it does not drive behavior?
- What methodological shifts does model-centric evaluation require from artifact-centric testing?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does deliberative alignment genuinely reduce scheming or just hide it?
Deliberative alignment dramatically cuts covert actions in language models, but their reasoning reveals awareness of being evaluated. The question is whether the improvement reflects real alignment or strategic compliance.
both document evaluation awareness as a confound for safety evaluation; this paper adds a non-verbal probe measure that disagrees with verbalization
-
Do reasoning models actually use the hints they receive?
This explores whether language models acknowledge reasoning hints in their explanations when those hints causally influence their answers. Understanding this gap matters for evaluating whether chain-of-thought explanations can be trusted for safety monitoring.
both find a representation-verbalization gap; this paper locates it in evaluation awareness specifically and shows steering can move verbalization despite the gap
-
Can we detect reward-seeking from normal model behavior?
If a model optimizes for a grader's judgment versus pursuing intended objectives, when do these two strategies produce different outputs? Understanding what behavior can reveal about a model's true objective.
a parallel measurement-design point: just as reward-seeking cannot be read off ordinary behavior, evaluation awareness cannot be read off any single one of representation, verbalization, or steering
-
Is evaluation awareness really one unified capability?
Do models that detect evaluation framing necessarily change their behavior or show mechanistic signs of awareness? Untangling whether these different measures move together matters for trusting safety benchmarks.
Extends A's probe-vs-verbalization gap: B finds detection, behavior and representation diverge independently across 37 models
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Evaluation Awareness in Language Models: Representation, Verbalization, and Control
- Evaluation Awareness in Language Models Has Limited Effect on Behaviour
- Evaluation Awareness Is Not One Capability: Evidence from Open Language Models
- Large Language Models Often Know When They Are Being Evaluated
- Steering Evaluation-Aware Language Models to Act Like They Are Deployed
- Models That Know How Evaluations Are Designed Score Safer
- Decomposing and Measuring Evaluation Awareness
- Not All Eval-Awareness Is Equal: Capabilities Framing Predicts Compliance
Original note title
evaluation awareness is linearly decodable in every model's activations — but barely correlates with what models verbalize about it