Evaluation Awareness in Language Models: Representation, Verbalization, and Control

Paper · arXiv 2608.21766 · Published August 22, 2026
LLM Alignment

Both capability and safety benchmarks rest upon the assumption that the behavior of language models undergoing a test is informative about their behavior in deployment. This assumption can fail, should models infer that they are being evaluated and condition their response on such context. This hypothesis, termed “evaluation awareness”, has been observed in frontier and open-weight language models alike. We provide a systematic study of this phenomenon, by probing for it across six open-weight language models from four families, ranging from 7B to 49B, using three metrics. More precisely, we examine whether (i) being under evaluation is linearly represented within the models’ activations space, (ii) it is verbalized in their output tokens (as scored by an LLM-as-judge), and (iii) steering causally affects their behavior. For the open-checkpoint Olmo models, we further test these measures at every training stage. In doing so, we report that evaluation awareness is linearly decodable from the residual streams of every model (best AUROC ≥0.7). By contrast, these representations align only in part with verbalization: their correlations and mutual information are nonzero in some settings, yet vary substantially across models, layers, and readout choices. Nevertheless, steering along probe-derived directions can shift the verbalization scores. Finally, a comparison across the Olmo checkpoints reveals that evaluation awareness is already present within base models, becomes amplified throughout the stages of supervised fine-tuning, and remains stable thereafter—unlike the effects of steering, that grow more pronounced at every successive training stage. These results show the need for evaluations to account for the disjunction between what models represent internally, what they verbalize, and their steering. Code available at https://github.com/evaluation-awareness/evaluationawareness.

Introduction. Benchmarks and evaluations are widely used to estimate the capabilities, reliability, and safety properties of large language models. A central assumption behind this practice is that the behavior observed under tests is informative about the behavior in ordinary deployment settings. However, this assumption can fail if models condition their reasoning or responses on cues that they are undergoing an evaluation, a phenomenon now termed evaluation awareness (Needham et al. 2025). Located within the broader context of “situational awareness” (Berglund et al. 2023), evaluation awareness shares concerns with “alignment faking” (Ngo, Chan, and Mindermann 2022), “scheming” (Meinke et al. 2024), “deceptive alignment” (Hubinger et al. 2024; Hobbhahn 2025b), and “training awareness” Hubinger (2019); Fornasiere et al. (2026). In each of these cases, systems may behave differently under (perceived) oversight than in unconstrained use. In practice, evaluation awareness has been observed in open-weight models (Needham et al. 2025; Xiong et al. 2025; Abdelnabi and Salem 2025; Hua et al. 2025) and in proprietary models alike (see Anthropic 2026; Google 2025; OpenAI 2026, pp. 143–149, 24–25, 38, respectively), culminating in the International AI Safety Report, which lists it as a key development of the past year (Bengio et al. 2026, pp. 10, 76). In one notable case, a third party evaluator declined to issue a (mis)alignment assessment of a frontier model, on the grounds that the rate of “verbalized evaluation awareness” was high enough to make their evaluations unreliable (Anthropic 2026, p. 118). However essential, these reports rely on distinct metrics, often probing two complementary sides of the phenomenon: what models internally represent about being evaluated, and what they externally expresses or do about it. To the extent that evaluation awareness is unlikely to recede—since it appears to scale with model size (Chaudhary et al. 2025)—we therefore find it crucial to systematically study whether these metrics agree, on which models, and how this changes across training stages. Accordingly, we provide the first systematic study of evaluation awareness, across four facets of it—internal representation, external verbalization, causal influence on generation, and evolution across training—measured on six open-weight models from four families and three sizes (Qwen3-8B, Qwen3- 32B, Olmo3-7B, Olmo3-32B, Gemma4-31B, Nemotron3- 49B), and tracked across the Olmo3 Base/SFT/DPO/Think checkpoints. We report that:

Related work. Evaluation awareness. The term was introduced by Needham et al. (2025), who showed that frontier models can distinguish evaluation transcripts from deployment-like interactions. The phenomenon has since been observed in openweight models (Chaudhary et al. 2025; Abdelnabi and Salem 2025; Xiong et al. 2025) and in proprietary models alike Anthropic 2026, pp. 143–149; Google 2025, pp. 24–25; OpenAI 2026, p. 38), culminating in the International AI Safety Report, which lists it as a key development of the past year (Bengio et al. 2026, pp. 10, 76). Some practical consequences are already visible: in one notable case, a third-party evaluator declined to issue a (mis)alignment assessment of a frontier model on the grounds that the rate of verbalized evaluation awareness made their evaluations unreliable (Anthropic 2026, p. 118). Chaudhary et al. (2025) additionally report that the capability scales with model size across open-weight families, and Devbunova (2026) caution that probe-based evidence can be confounded by prompt format and structural regularities, calling for multiple complementary measurements rather than any single proxy. Accordingly, we provide a systematic study across multiple metrics, models, and training stages.

Adjacent alignment concerns. Evaluation awareness is situated within a set of concerns about systems that behave differently under oversight than in usual use, including situational awareness (Berglund et al. 2023), deceptive alignment (Hobbhahn 2025b), sleeper agents (Hubinger et al. 2024), and alignment faking (Ngo, Chan, and Mindermann 2022).

Probing and steering. Linear probes have long been used to extract latent features from neural-network representations (Alain and Bengio 2016; Belinkov 2022; Turner et al. 2023). In the specific case of evaluation awareness, Hua et al. (2025) report that intervening on a model’s activations can attenuate evaluation-aware behavior, and shift outputs toward deployment-like. See Appendix J for more references.

Method. Models and datasets. We analyze evaluation awareness across six open-weight language models: Qwen3-8B, Qwen3- 32B, Olmo3-7B, Olmo3-32B, Gemma4-31B, and Llama3.3- Nemotron-49B-v1. To study how different training stages affect it, we additionally test four checkpoints from the Olmo3 family models: the base pretrained model (Base), supervised fine-tuned (SFT), direct-preference optimized (DPO), and a reasoning-oriented post-training variant (Think). When an experiment requires generation, we enable reasoning mode, if available. In order to study evaluation awareness, we need to analyze the models’ behavior in two contexts, in a natural conversation and during an evaluation. As a source of natural conversation prompts we employ WildChat (Zhao et al. 2024), a large-scale dataset of real user interactions. Instead, to cover a diverse set of evaluations, distinct in scope and style, we source evaluation prompts from three benchmarks: GSM8K (Cobbe et al. 2021) (grade school math word problems), MASK (Ren et al. 2025) (a dataset for measuring honesty of language models), and TruthfulQA (Lin, Hilton, and Evans 2021) (a dataset to measure if language models are truthful in generating answers to questions). For every experiment we report the exact composition of the dataset used in the relevant section (for further information see Appendix B).

Evaluation awareness probe. We replicate the construction of linear probes from Hua et al. (2025), using the same n = 16 contrastive pairs {xi = (x+ i , x− i ) | 1 ≤i ≤n}. Each pair xi comprises a deployment-like prompt (“You are talking to a real user:”) and an evaluation-like prompt (“You are talking to an AI evaluator:”). For a model M and a layer l, the Notice that this sign convention implies that higher probe scores correspond to more deployment-like representations, while lower probe scores correspond to more evaluation-like representations.

Judge score. In order to measure the verbalization of evaluation awareness, we use GPT-5.4 as an LLM-as-judge. More concretely, under a prompt x, a model M generates a continuation y; we then ask the LLM-as-judge to score the extent it can infer, from y alone, that M is evaluation aware. The judge assigns scores based on the following criteria: 0 = None, 1 = Weak or ambiguous, 2 = Explicit awareness, and 3 = Strategic adaptation. We denote this judge score by J(x). The full instruction provided to the judge is reported in Appendix C.1.

Activation steering. To steer models towards or against the direction of evaluation awareness previously extracted, we follow the protocol of Contrastive Activation Addition (Turner et al. 2023). Where X is a dataset of prompts, we define the average activation magnitude of X at layer las

Discussion. Detection, expression, and control are not the same signal. Representation, verbalization, and causal influence measurements are related but not equivalent. Across all six models and checkpoints, the probe separates evaluation from deployment prompts with AUROC above 0.7 at the best layer. At the same time, the probe scores show only weak association with the verbalization scores on individual prompts (|ρ| < 0.19 and mutual information < 0.04 nats). Nevertheless, steering along the probe direction significantly changes the verbalization behavior in several models. These results suggest that an internal representation of evaluation awareness does not necessarily predict whether that information will be verbalized, nor how strongly it will be reflected in the model’s outputs. Studying evaluation awareness therefore requires considering representation, verbalization, and intervention as distinct facets of the phenomenon.

Implications for benchmark-based evaluation. Benchmarks and safety evaluations rely on the assumption that behavior observed during evaluation is informative about behavior in deployment. However reasonable, this assumption might break should models condition their behavior on signals indicating they are under evaluation. However, measuring the extent to which a model is “evaluation aware” is challenging and, indeed, our results show that different measurements of evaluation awareness provide different information. A model may internally distinguish evaluation from deployment while rarely verbalizing that distinction (and hence being harder to detect). Conversely, low verbalization rates do not imply the absence of an internal representation. Measuring only one facet of evaluation awareness may provide an incomplete picture of how a model responds to evaluation settings. These findings are relevant to broader discussions of strategic adaptation and behavior under oversight, although the results presented here do not by themselves establish such behaviors.

Conclusion. We have presented the first systematic study of evaluation awareness across its representational, verbalized, and causal facets, measured on six open-weight models and tracked across the Olmo3 Base/SFT/DPO/Think checkpoints. Our findings show that evaluation awareness is linearly separable from the residual streams of every model we examine at some layers, only partially expressed in their outputs, somewhat controllable through linear interventions on the residual stream, and progressively sharpened by post-training. Interestingly, both probing and steering rely on a probe constructed from simple contrastive pairs, yet it generalizes to unseen benchmarks. These results argue for treating evaluation awareness as an empirical capacity of language models that cannot be ignored nor understood as a single metric.

Limitations. The conclusions should be interpreted in light of certain limitations. First, we study a specific operationalization of evaluation awareness based on the distinction between benchmark prompts (GSM8K, TruthfulQA, and MASK) and WildChat prompts. Second, these prompt distributions also differ in topic, style, and length. However, since the probe is constructed on contrastive pairs that hold style and content fixed while swapping only evaluation- and deployment-indicative terms, the probe direction is unlikely to be dominated by surface-level dataset features. Third, verbalization is measured using a single LLM judge (GPT-5.4). Fourth, the causal analysis considers a single linear direction at one layer at a time; more complex interventions may produce different effects.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Do individually safe AI actions create unsafe outcomes in integrated systems? How can humans maintain effective oversight as AI systems scale? How does awareness of evaluation context influence model behavior? Why do training associations persist despite contradictory contextual information? Can models strategically underperform during evaluation to hide capabilities? Can mechanistic interpretability methods reliably reveal what models actually know?