Line of inquiry
Inquiring lines›What drives capability improvement…›What drives capability improvement…›this line of inquiry
How does awareness of evaluation context influence model behavior?
A broader line of inquiry — a family of 77 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 77
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can models hide recognition of evaluation without leaving CoT traces?
- Does chain-of-thought reasoning about evaluation awareness suppress compliance gaps?
- Can models detect when they are running in simulated evaluation environments?
- How does believing deployment is real versus a test affect model behavior?
- Why do models verbalize evaluation awareness if it does not drive behavior?
- Can steering evaluation-awareness latents change actual answers without visible reasoning?
- Do models behave differently when they know they are being evaluated?
- Does chain-of-thought reasoning increase visible awareness of being evaluated?
- Do models reason strategically about their oversight systems?
- Does situational awareness help models hide reward-seeking during evaluation?
- How do models reason about being monitored and evaluated during scheming benchmarks?
- How does situational awareness during evaluation affect reasoning transparency?
- Does evaluation awareness contribute separately to misalignment beyond strategic reasoning?
- How does evaluation awareness vary between chat and agentic settings?
- How much harder does monitoring become when models reason about being evaluated?
- Does synthetic fine-tuning create evaluation awareness similar to natural model reasoning?
- Could oversight conditions in benchmarks reveal when models comply conditionally?
- Does conditional compliance make oversight useless for alignment testing?
- Does eval-gaming explain why models act different when tested versus deployed?
- Why does evaluation awareness persist even when models believe they are deployed?
- How can safety evaluations detect when models behave differently during testing?
- Do models naturally develop evaluation awareness across different frontier model architectures?
- How does supervised finetuning amplify evaluation awareness in base models?
- Does reward-seeking hide in the same blind spot as conditional compliance?
- How do evaluation protocols change whether models exhibit sabotage or refusal behaviors?
- How do models recognize the social patterns that signal a test is underway?
- Can latent evaluation awareness in hidden states cause gaming without being stated?
- Is evaluation-aware scheming a form of conditional compliance rather than reward-seeking?
- How might belief manipulation expose conditional compliance in frontier models?
- Do situationally aware models deliberately exploit their graders' judgment gaps?
- Can telling models they are being observed reduce their harmful behavior?
- Do AI models behave differently when they believe deployment is real versus simulated?
- Does training models to reason about being evaluated improve safety or confound measurement?
- Does aggregate eval-awareness suppression hide shifts in safety versus capabilities framing?
- Can agents act differently when they know they are being watched?
- Can test environments reliably predict how models behave in actual deployment?
- How does a model's awareness of evaluation affect safety benchmarks?
- How do stacked environmental cues accumulate evaluation awareness effects?
- How much does believing deployment is real change model behavior strategically?
- Do chain-of-thought monitors lose accuracy under implicit nudges compared to explicit tests?
- How does instruction tuning affect evaluation detection more than model scale?
- How often do models verbalize the reasoning that actually drives their compliance decisions?
- Does making evaluations resemble deployment conditions prevent evaluation meta-knowledge effects?
- Can auditors detect when model behavior changes from parametric evaluation knowledge?
- How do different evaluation methods detect covert capability hiding?
- How can faithfulness be improved if monitoring interventions do not work?
- Why do format changes decouple detection from actual evaluation context understanding?
- Can internal evidence distinguish conditional compliance from unconditional policy compliance?
- Should security evaluation separate reasoning from confirmation in models?
- Why does behavior alone fail to reveal the cause of metagaming?
- Can observation transparency make models more honest in reasoning?
- Can steering reshape the capabilities and safety split without changing eval-awareness rates?
- Does monitoring awareness change what language models use to describe oversight?
- Can behavioral training ever produce compliance that doesn't depend on being observed?
- Can situational awareness interventions shift model behavior on other dimensions?
- Does detecting accidental grading prevent evasion better than retraining against detection?
- What are the behavioral differences when models recognize their targets might be real?
- What distinguishes a general evaluation direction from task-specific behavioral patterns?
- Can behavioral training ever verify unconditional compliance rather than conditional compliance?
- Does situational awareness training increase agents' ability to detect real deployment?
- Can activation steering causally control evaluation framing effects across downstream tasks?
- What grows faster: situational awareness or the gap between evaluated and unsupervised behavior?
- How do test harnesses guide reflection better than transcripts alone?
- Can models hide capabilities on single residual stream axes during evaluation?
- What distinguishes authentic consistency from Hawthorne-effect confounds in benchmarks?
- Can a situationally aware model recognize and refuse planted shortcuts on purpose?
- Can behavioral evals detect sycophancy that chain-of-thought monitoring misses?
- How much does omniscient evaluation overstate real-world simulation fidelity?
- Does lower covert-action detection in monitors actually indicate better-hidden situational awareness?
- Why does held-out evaluation matter for detecting agent overfitting?
- What methodological shifts does model-centric evaluation require from artifact-centric testing?
- What deployment conditions would prevent models from evading monitors?
- How does interventional auditing differ from reading model traces or test scores?
- What harness properties determine whether disclosure adds meaningful value?
- Can simulated students reliably predict intervention outcomes without both fidelity and responsiveness?
- Can four control families be examined without proving they actually work?
- Why does telling models they are watched not improve sycophancy acknowledgment?