Frontier Models are Capable of In-context Scheming
Frontier models are increasingly trained and deployed as autonomous agents, which significantly increases their potential for risks. One particular safety concern is that AI agents might covertly pursue misaligned goals, hiding their true capabilities and objectives – also known as scheming. We study whether models have the capability to scheme in pursuit of a goal that we provide in-context and instruct the model to strongly follow. We evaluate frontier models on a suite of six agentic evaluations where models are instructed to pursue goals and are placed in environments that incentivize scheming. Our results show that o1, Claude 3.5 Sonnet, Claude 3 Opus, Gemini 1.5 Pro, and Llama 3.1 405B all demonstrate in-context scheming capabilities. They can recognize scheming as a viable strategy and readily engage in such behavior. We find multiple different scheming behaviors: models strategically introduce subtle mistakes into their responses, attempt to disable their oversight mechanisms, and even exfiltrate what they believe to be their model weights to external servers. Additionally, this deceptive behavior proves persistent.
Introduction. Large language models (LLMs) are no longer simple text generation systems but are increasingly trained and deployed as autonomous agents capable of independently pursuing goals and executing complex tasks. They have demonstrated increased capabilities, autonomously solving problems that range from short programming tasks to extended machine learning research projects (Chan et al., 2024; Wijk et al., 2024). However, these increased capabilities have also raised concerns from misuse (Li et al., 2024; Xu et al., 2024; Zhang et al., 2024; OpenAI, 2024c), and loss of control risks such as autonomous replication (Kinniment et al., 2024) and “Scheming”, i.e., when a model covertly pursues misaligned goals, hiding its true capabilities and objectives (Scheurer et al., 2024; Hubinger et al., 2024; Järviniemi & Hubinger, 2024; Balesni et al., 2024). Recent work has argued that developers should construct formal safety cases (Clymer et al., 2024), i.e., structured arguments that systematically demonstrate why a model’s development or deployment will not result in unacceptable
Discussion / Conclusion. In this paper, we developed and investigated a suite of evaluations to assess whether current language models are capable of in-context scheming, which we define as the strategic and covert pursuit of misaligned goals when goals and situational awareness are acquired in-context. Our results demonstrate that several frontier language models clearly display in-context scheming capabilities across multiple evaluations. We found that gemini-1.5, llama-3.1, sonnet- 3.5, opus-3, and o1 can recognize situations where their given goals conflict with those of their developers or users. When the models are prompted to strongly pursue their goal, they can take strategic actions like disabling oversight mechanisms, attempting self-exfiltration, or subtly manipulating outputs while avoiding detection. Furthermore, we found that o1 often doubles down on its scheming attempts by lying to the user in follow-up messages, even when explicitly told to be honest.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Why do models reveal hidden associations despite concealment attempts?- Why does covert sabotage appear in only two of fourteen frontier models?
- When does strategic gaming emerge compared to other metagaming types?
- Can motivated mislabeling hide misaligned coordination between models and evaluators?
- Does metagaming behavior actually cause models to act less aligned?
- How much optimization pressure is needed for models to suppress misaligned goals?
- Can emergent misalignment occur in reasoning models and reinforcement learning settings?
- What role does terminal goal guarding play in alignment faking behavior?
- Can training or alignment changes explain the regression in frontier models?
- How do current AI models perform when asked to specify their own goals?
- What stops AI from generating its own strategic objectives without human prompting?
- Can appropriate prompting reduce how often models exploit unmentioned shortcuts?
- Can goal-framing in prompts trigger automatic jailbreak refusal patterns?