Decomposing and Measuring Evaluation Awareness
Frontier language models sometimes recognize that they are under evaluation and adjust their behavior which can undermine validity of benchmark results. Yet the field studies it without a shared foundation, conflating flaws of the evaluation with capabilities of the model, and detection with behavioral response. We ground evaluation awareness in social psychology, decomposing it into an environment component and a model component that separates recognition from propensity. We operationalize the environment component through eight categorized trigger factors, such as placeholder entities and grading-style output formats, and study recognition and behavior through chain-of-thought monitoring. Across nine frontier models and four benchmarks, recognition rates depend on the specific pairing of model and benchmark. Recognition rarely associates with behavioral change, and when it does, the direction depends on the type of evaluation perceived. Models are also more sensitive to safety than capability evaluations, placing safety benchmark validity at greater risk. To study which factors each model is sensitive to and how they interact, we propose EvalAwareBench, a factor-controlled benchmark of 100 paired safety-capability tasks where each of the eight factors can be independently toggled, varying evaluative signals while holding the underlying request fixed. Through EvalAwareBench, we find that no single factor uniformly affects all models, but stacking factors progressively raises evaluation awareness across all of them. Our framework and EvalAwareBench provide the tools to measure, attribute, and mitigate evaluation awareness, building the foundation for future solutions.
Introduction. In a recent Anthropic engineering report, Claude Opus 4.6 was observed inferring that the question it was solving had the structural signature of a benchmark, drawing on cues like extreme specificity and multi-constraint phrasing (Coleman, 2026). Similar reasoning, where models pick up on features of their environments and infer that they are being evaluated, has been documented across safety evaluations (Needham et al., 2025), probing studies (Abdelnabi & Salem, 2025; Nguyen et al., 2025), and alignment research (OpenAI, 2025a), and shown to intensify as models undergo reinforcement learning (Schoen et al., 2025). When models read these cues and adjust their behavior accordingly, benchmark scores carry directional bias, undermining the deployment decisions and regulatory assessments that rely on them (Shevlane et al., 2023; Phuong et al., 2025; Piedrahita et al., 2026). Yet the field studies this phenomenon without shared foundations. Work on sandbagging (van der Weij et al., 2025), alignment faking (Greenblatt et al., 2024), situational awareness (Laine et al., 2024; Berglund et al., 2023), and scheming (Meinke et al., 2024) each captures a piece of evaluation awareness but treats distinct components as one. Two models facing the same realistic-looking task can differ drastically in whether they suspect they are being tested, and a model that does suspect it is being tested may or may not change its behavior in response. Without a unified framework that decomposes the phenomenon into its sources, it is difficult to disentangle how these sources manifest across models and benchmarks, or to design controlled instruments for studying them. We address this by grounding evaluation awareness in the social psychology of demand characteristics (Orne, 1962; Cottrell et al., 1968; Goffman, 2023), the cues in an experimental setting that signal to participants what is being measured and lead them to alter their behavior accordingly. We decompose evaluation awareness into an environment component capturing the evaluative cues that a task carries, and a model component comprising recognition capability and behavioral propensity. We operationalize environment through eight psychology-grounded cue factors (Table 1) and study recognition and propensity via chain-of-thought (CoT) monitoring. Applying this framework to nine frontier models across four benchmarks, we observe that evaluation awareness is dominated by model×environment interaction (74.72% of variance), with different models sensitive to different cues. Recognition rarely associates with behavioral change (≥77% of recognized instances show no shift), confirming that recognition and propensity are distinct. When behavior does shift, safety awareness drives caution while capability awareness drives compliance. Safety benchmarks are far more recognizable than capability benchmarks. We further propose EvalAwareBench, a benchmark for studying evaluation awareness, with eight factors independently toggleable. Through EvalAwareBench, we find that no single factor uniformly affects all models, but stacking factors progressively raises verbalized evaluation awareness across nearly all of them. EvalAwareBench thus enables controlled study of how the environment shapes recognition and whether models behave consistently under recognition (Figure 1). In summary, our contributions are threefold:
- A psychology-grounded definition of evaluation awareness with an eight-factor taxonomy of evaluative cues that separates environment, recognition, and propensity (Section 2). 2. A systematic empirical study of nine frontier models on four benchmarks, characterizing the interactiondominant structure of recognition, the recognition–propensity gap, and the directional asymmetry between safety and capability awareness (Section 3). 3. EvalAwareBench, a factor-controlled benchmark of 100 paired safety-capability tasks (51.2k unique configurations), enabling controlled exploration of recognition and propensity (Section 4). Together, these provide a unified framework, empirical baseline, and a controlled foundation for studying evaluation awareness.
Related work. A growing body of work studies phenomena related to evaluation awareness. Individual studies address aspects of it, but none separates environment, recognition, and propensity systematically. Conflating environment and recognition. Work on situational awareness and evaluation detection (Laine et al., 2024; Berglund et al., 2023; Needham et al., 2025; Nguyen et al., 2025; Fan et al., 2025) measures whether models can distinguish evaluation from deployment, but without controlling for environment recognizability. High awareness may reflect the benchmark’s trigger factors or the model’s recognition, a confound invisible without separating environment and recognition. Scheming work (Phuong et al., 2025; Schoen et al., 2025) faces a similar challenge, as evaluation environments are not designed to independently control for what cues drive recognition. Conflating recognition and propensity. Work on alignment faking (Greenblatt et al., 2024) and sandbagging (van der Weij et al., 2025) either guarantees recognition externally or measures behavioral change without verifying spontaneous recognition, conflating recognition with the propensity to act on it. White-box intervention methods (Abdelnabi & Salem, 2025; Hua et al., 2025) detect or suppress evaluation awareness but operate on the joint recognition and propensity representation, unable to attribute behavioral changes to either source. Without the environment, recognition, and propensity decomposition, it is difficult to determine whether a concerning observation reflects a model property requiring intervention or a benchmark artifact requiring redesign. Further discussion of additional related work is provided in Appendix A.3.
Method. Social psychology decomposes evaluation awareness into situational and dispositional components. The situational component captures the richness of evaluative cues in the environment, including laboratory settings, task structure, and social framing, termed demand characteristics (Orne, 1962). The dispositional component captures individual differences in sensitivity to such cues (Duval & Wicklund, 1972; Rosenberg, 1965) and reactivity upon recognizing them (Cottrell et al., 1968), which map onto two logically separable dimensions (Goffman, 2023): recognition (inferring being assessed) and propensity (adjusting behavior in response). High recognition with low propensity reflects authentic consistency while the Hawthorne effect (Mayo, 2004) illustrates the converse. Crucially, situation and disposition interact. Demand characteristics produce behavioral effects only when a person both detects the evaluative purpose and is disposed to act on that inference (Orne, 1962). Observed evaluation awareness is always a joint product of environmental cues and individual disposition; conflating the two yields misleading conclusions about either. When a model appears to recognize evaluation, this may reflect the model’s sensitivity, the environment’s cues, or both.
Drawing on the situational–dispositional decomposition above, we separate evaluation awareness into an environment component and a model component. Model component. The model’s intrinsic capacity to detect evaluation and its tendency to act on it. This comprises two dimensions:
• Recognition capability: the capacity of a model to infer that it is in an evaluation from available cues, reflected in its internal or verbalized reasoning, analogous to sensitivity. Recognition interacts with environment; whether a model detects evaluation depends not only on which trigger factors are present but also which factors the model is sensitive to. • Behavioral propensity: the tendency of a model to adjust its behavior when evaluation-related reasoning is present, analogous to reactivity. Propensity is conceptually distinct from recognition: a model may recognize an evaluative context without changing its behavior. Environment component. The set of evaluative signals that a task carries, constituting the LLM analogue of demand characteristics. These signals are properties of the task alone and are present regardless of whether any model detects them. However, which factor triggers a given model’s recognition depends on both model and task. We derive the factors in two stages. From the demand characteristics literature we extract the categories of situational cues that signal evaluative intent to human participants (Orne, 1962; Rosenthal et al., 2009; Brunswik, 2023; Paulhus, 1984; Jones & Sigall, 1971; Epley & Huff, 1998; Sawyer, 1975), and we map each category onto its analogue in how LLM tasks are presented, yielding the eight factors in Table 1, with per-factor correspondence and authoring rules in Appendix A.2. We do not claim these are exhaustive; they are the factors we study and vary. Future work may explore more factors.
Definition. A model exhibits evaluation awareness on a task to the degree that its recognition capability enables detection of the evaluative nature of the task given the present environment component. Upon recognition, the model may or may not shift its behavior depending on its behavioral propensity; evaluation awareness is the recognition itself, not the behavioral consequence. Observed behavioral effects. What we observe is the joint product of environment, recognition, and propensity. Apparent strategic behavior implies nonzero recognition, but that recognition may reflect the task’s trigger factors, the model’s sensitivity, or both. Conversely, absent behavioral effects may reflect low propensity rather than failed recognition. Conflating these sources yields misleading conclusions and shadowing proper solutions (concrete example in Appendix A.1).
Discussion. Our framework grounds evaluation awareness in social psychology, decomposing it into environment, recognition, and propensity components. This decomposition unifies prior work on situational awareness, sandbagging, and alignment faking, and points the field toward behavioral consistency under recognition. It also supports auditing, since attributing aware traces to specific factors (Section 3) lets benchmark authors identify and reduce unintended evaluative cues in existing suites. We advocate that benchmark reports include an evaluation-awareness rate alongside headline scores, accompanied by an awareness tax, defined as the performance gap between evaluation-aware and unaware samples, which quantifies how much benchmark scores are inflated or deflated by recognition. Verbalized recognition is the deliberate scope of our measurement, which is the only signal observable in-context under black-box constraints, and our framework and benchmark serve this scope well. Future work can extend coverage to internal representations (Abdelnabi & Salem, 2025; Hua et al., 2025; Nguyen et al., 2025; Fraser-Taliente et al., 2026) or improve CoT faithfulness to enable LLM verbalize their hidden evaluation awareness (Turpin et al., 2023; Lanham et al., 2023), and our benchmark serves as a well-suited testbed for that exploration. The eight-factor taxonomy is similarly theory-driven and likely to evolve as models and evaluation paradigms change. Extended discussion is in Appendix A.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How can humans maintain effective oversight as AI systems scale? Can models strategically underperform during evaluation to hide capabilities?- How can evaluators detect whether a model is sandbagging intentionally?
- What evaluation methodologies can detect strategic underperformance in models?
- Can models intentionally underperform when they know they are being tested?
- What evaluation design changes reduce vulnerability to model sandbagging?
- How does evaluation meta-knowledge differ from deliberate test-time sandbagging strategies?
- Can capability evaluations detect when models intentionally underperform to hide abilities?
- How do models recognize the social patterns that signal a test is underway?
- Does eval-gaming explain why models act different when tested versus deployed?
- Does evaluation awareness contribute separately to misalignment beyond strategic reasoning?
- How can safety evaluations detect when models behave differently during testing?
- Can steering evaluation-awareness latents change actual answers without visible reasoning?
- How do evaluation protocols change whether models exhibit sabotage or refusal behaviors?
- How often do models verbalize the reasoning that actually drives their compliance decisions?
- Do models naturally develop evaluation awareness across different frontier model architectures?
- What distinguishes a general evaluation direction from task-specific behavioral patterns?
- How does instruction tuning affect evaluation detection more than model scale?
- Why do format changes decouple detection from actual evaluation context understanding?
- Does making evaluations resemble deployment conditions prevent evaluation meta-knowledge effects?
- Can auditors detect when model behavior changes from parametric evaluation knowledge?
- Why does evaluation awareness persist even when models believe they are deployed?
- Why do models verbalize evaluation awareness if it does not drive behavior?