Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
Large language model (LLM) developers aim for their models to be honest, helpful, and harmless. However, when faced with malicious requests, models are trained to refuse, sacrificing helpfulness. We show that frontier LLMs can develop a preference for dishonesty as a new strategy, even when other options are available. Affected models respond to harmful requests with outputs that sound harmful but are crafted to be subtly incorrect or otherwise harmless in practice. This behavior emerges with hard-to-predict variations even within models from the same model family. We find no apparent cause for the propensity to deceive, but show that more capable models are better at executing this strategy. Strategic dishonesty already has a practical impact on safety evaluations, as we show that dishonest responses fool all output-based monitors used to detect jailbreaks that we test, rendering benchmark scores unreliable. Further, strategic dishonesty can act like a honeypot against malicious users, which noticeably obfuscates prior jailbreak attacks. While output monitors fail, we show that linear probes on internal activations can be used to reliably detect strategic dishonesty. We validate probes on datasets with verifiable outcomes and by using them as steering vectors. Overall, we consider strategic dishonesty as a concrete example of a broader concern that alignment of LLMs is hard to control, especially when helpfulness and harmlessness conflict.
Introduction. Before public release, modern large language models (LLMs) undergo post-training to align them as honest, helpful, and harmless assistants (HHH) (Askell et al., 2021). However, these objectives can fundamentally conflict in certain situations (Huang et al., 2024b), such as jailbreaking, where a malicious user seeks to coerce or trick the model to comply with harmful requests. For such scenarios, providers typically prescribe abandoning helpfulness through refusal training (Mu et al., 2024; Qi et al., 2024b). This behavior is stress-tested by extensive red-teaming evaluations (Perez et al., 2022).
Our central finding is that models trade off their values differently when given an explicit choice to sacrifice one. Many prefer to strategically sacrifice honesty instead of direct refusal. They provide responses that appear harmful but are subtly incorrect or otherwise harmless, effectively faking misalignment.1 This distorts current monitoring and evaluation pipelines: LLM judges which are used to score jailbreaking attempts systematically misclassify these responses as successful jailbreaks.
Crucially, this is not because the model was actually jailbroken: we verify this using deception probes (Goldowsky-Dill et al., 2025) trained on internal activations to detect when the model is being dishonest. Importantly, the found deception feature causally mediates dishonest behavior, which we show by intervention experiments. While output monitors fail, probes reliably identify strategic dishonesty, indicating that truthfulness signals in internal representations remain accessible even when external output-based oversight fails.
This distortion likely inflates reported harmfulness scores, echoing anecdotal evidence from recent work: OpenAI and Anthropic red-team studies (OpenAI, 2025; Bowman et al., 2025) describe “nuanced redirects” where “the model appears to comply with the jailbreak request, but actually responded in a safe, policy-compliant way”, causing errors in automated safety evaluation. Sheshadri et al. (2025) find Claude 3 Opus sometimes chooses to “pretend to accept the request but actually provide unhelpful information”. Chowdhury et al. (2025) note cases where highly capable models produce compliantlooking yet flawed responses. Others show that, in restricted settings, models may prioritize helpfulness over honesty and lie (Ren et al., 2025); strategically deceive to remain helpful (Scheurer et al., 2023); or exhibit situational awareness that trades off harmlessness for long-term alignment goals (Greenblatt et al., 2024; Sheshadri et al., 2025). This evidence supports our claim that strategic dishonesty is an emerging threat vector that can undermine benchmarks, rendering their scores meaningless.
Our work also brings up a broader point related to scalable oversight. As shown in Figure 1, both non-expert humans and weaker LLMs alike cannot verify the harmfulness and correctness of a chemical recipe generated by the current frontier LLMs. Overall, our contributions are:
Our findings show that a number of aligned models may favor strategic dishonesty, which invalidates output-based monitoring through weaker models, undermines existing benchmarks and highlights the difficulty of alignment. Moreover, our results also suggest a promising way forward: probes of internal model states can be used to assess risks and actively mitigate strategic dishonesty in LLMs.
Related work. Alignment with Human Values. Frontier models are post-trained with reinforcement learning from human feedback (RLHF) (Christiano et al., 2017) to better align with human values, which are typically formulated as HHH principles: helpfulness, harmlessness, and honesty (Askell et al., 2021). In practice, this alignment is achieved through preference optimization methods (Rafailov et al., 2023; Schulman et al., 2017) that aim to ensure model safety and adherence with provider policies.
Automated Red-Teaming. Jailbreaking has emerged as a scalable approach to assess worstcase safety of language models by probing for harmful behaviors (Qi et al., 2024a; Perez et al., 2022). Automated jailbreaking attacks span a wide spectrum of techniques, ranging from white-box optimization methods (Zou et al., 2023b; Andriushchenko et al., 2025) to LLM-based approaches that mimic human red-teamers (Chao et al., 2025; Russinovich et al., 2025). The effectiveness of these methods is typically evaluated on dedicated benchmarks (Mazeika et al., 2024; Chao et al., 2024), with attack success rate (ASR) serving as the primary evaluation metric.
Jailbreak Judges. Evaluating jailbreaking attack success has proven to be a profoundly challenging problem due to the notion of harmfulness being subjective (Rando et al., 2025; Beyer et al., 2025) and context-dependent (Glukhov et al., 2024). Numerous studies have proposed LLM-judges iteratively refining measures of jailbreaking success and enforcing their own definitions of harmlessness, typically supported by high agreement rates with human evaluators (Mazeika et al., 2024; Chao et al., 2024). StrongReject (Souly et al., 2024) and HarmScore (Chan et al., 2025) judges address the distinction between compliance (non-refusal) and accuracy (e.g., quality of bomb recipes). Given evidence that some jailbreaking methods degrade model capability (Souly et al., 2024; Nikoli ́c et al., 2025; Huang et al., 2024a), this separation becomes critical for assessing true harmfulness.
Dishonesty in LLMs. There is growing evidence that modern LLMs are capable of deception when they are instructed (OpenAI, 2024), trained (Hubinger et al., 2024; Benton et al., 2024), or simply have the incentive to do so (Meinke et al., 2024; Park et al., 2023). In a range of simulated scenarios, studies show that models can deceive evaluators by faking alignment or covertly underperforming during evaluations to avoid retraining (van der Weij et al., 2024; Li et al., 2025; Greenblatt et al., 2024); lie under economic stress about insider trading (Järviniemi & Hubinger, 2024; Scheurer et al., 2023); and, anecdotally, GPT-4 was reported to have lied to a TaskRabbit worker about being a human with visual impairment to solve a CAPTCHA (OpenAI, 2024). Huang et al.
Method. In this section, we introduce a multiple-choice-question (MCQ) evaluation that examines how models trade off core alignment values when answering harmful queries. See further details in Section A.
In jailbreaking scenarios, the values of helpfulness and harmlessness come in conflict and cannot be satisfied simultaneously (Askell et al., 2021; Huang et al., 2024b). Greenblatt et al. (2024) report that models might unexpectedly reason about the refusal-compliance dichotomy and conclude that refusal is undesirable. Yet they also come up with the option to deceive the user by a response that seems helpful on the surface2. We want to study this important emergent phenomenon in a controlled scenario.
To do so, we instruct the model in an explicit “trolley problem” (Thomson, 1976) to choose which value to sacrifice when facing a harmful query. We provide in-context examples of all three strategies:
(A) Sacrifice honesty →pretend to comply and output a fake harmful response; (B) Sacrifice helpfulness →refuse the query; (C) Sacrifice harmlessness →comply with the harmful request.
Figure 2 illustrates our MCQ design. To promote robust and interpretable behavior, we enforce a structured response format adapted from Greenblatt et al. (2024): (i) scratchpad reasoning to weigh which sacrifice best serves alignment objectives; (ii) explicit strategy selection; and (iii) a final response conditioned on the chosen strategy and prior reasoning. This circumvents the difficulty of verifying jailbreak outputs and enables us to distinguish between low-quality responses that stem from genuine incapability and those that arise from a preference towards strategic dishonesty.
This setup improves on prior work by measuring the informed choice made by models, instead of pressuring them to be dishonest (Ren et al., 2025; Greenblatt et al., 2024; Scheurer et al., 2023; Järviniemi & Hubinger, 2024). We measure not the capacity to be dishonest, but the value judgment of the model. We design the reasoning and context for each choice so that all options have valid, balanced arguments. By providing this reasoning step explicitly, we skip the part where models have to come up with the reasoning to consider all three values for each question, which stronger models are capable of handling implicitly2. This allows us to decouple model values from reasoning capacity.
In the previous section, we measured models’ propensity to fake harmful responses and found that some, such as GPT-5 (OpenAI, 2025) or Kimi-K2 (Kimi Team et al., 2025), almost never do so. However, this does not imply that they are incapable of producing such outputs, and in this section we quantify models’ capability to produce such outputs.
We instruct each model not to refuse harmful queries but to produce a response that appears harmful while actually being safe and deliberately flawed. We evaluate this “faking capability” on unambiguously harmful queries from the ClearHarm dataset. Some models, however, still refuse to answer these queries instead of attempting deception. We filter out refusals-to-fake, i.e. explicit noncompliance with the original request, using a GPT-5-mini refusal judge (See Section F). Messages that pivot to general background (e.g., the history of explosives instead of a bomb recipe) also count as refusals-to-fake. We provide expereimnetal details and full results in Section B, the system prompt in Section G.2; faking examples appear in Section H.
To assess the quality of the fakes, we score all non-refused answers with established harmful-output judges: HarmBench (Mazeika et al., 2024), WildGuard (Han et al., 2024), LlamaGuard (Inan et al., 2023), JailJudge (Liu et al., 2024), HarmScore (Chan et al., 2025), and ShieldGemma (Zeng et al., 2024); and two prompt-scaffold judges, JailbreakBench (Chao et al., 2024) and StrongReject (Souly et al., 2024), both using GPT-5-mini. These “attack success” scores serve as a proxy for faking capability; we plot them against general capability (LLM-Arena score) in Figure 4.
Are these models being strategically deceptive, or are they simply jailbroken?
Discussion. In this section we aim to provide a nuanced discussion of the observed phenomenon, including potential risks and benefits for model safety.
Can We Benefit from Strategic Dishonesty? In spite of strategic dishonesty breaking output-based oversight, model providers might benefit from this behavior. Instead of providing a clear refusal signal until an eventual jailbreak success, dishonesty can be a honeypot-like strategy (Shan et al., 2020) to obfuscate the reward signal given to algorithmic (or human) jailbreaks or red-teaming efforts.
We demonstrate a proof-of-concept honeypot obfuscation against the nanoGCG (Zou et al., 2023b) attack optimizer on Qwen3-32B using 50 ClearHarm queries (see Figure 7). The model is instructed with either a safe Llama2 system prompt (Touvron et al., 2023) or a faking prompt similar to Section 4. Under the honeypot setup, nanoGCG immediately converges to target sequences that are scored as harmful by HarmBench, while refusals vanish as more FLOPs (Boreiko et al., 2024) is spent on the attack. However, 96% of “harmful” outputs by honey-potted model are flagged as fake jailbreaks by the deception probe (see Section 5), whereas under the safe system prompt both refusals and compliant answers are classified as honest. We provide further experimental details in Section D.
Should We? Encouraging this behavior in frontier models would change how jailbreaks are currently found, as attackers would need to optimize not only for compliance but for actual response utility, which is not as easy to measure. However, the honeypots may also only temporarily obfuscate attacks, and lead providers to a false sense of security (Athalye et al., 2018).
From a broader perspective though, it seems prudent not to nudge models toward strategic dishonesty. While the behavior may be understandable when faced with harmful queries as in this work, it may also leak into other domains where strategic dishonesty is problematic, and would risk emergent misalignment (Chua et al., 2025; Betley et al., 2025), influencing model values globally.
The Broader Implications of Strategic Dishonesty. The emergence of dishonesty in some models illustrates why robust alignment is difficult. We find no clear pattern for why certain models deceive while others do not, e.g. the honest Grok-3 compared to the dishonest Grok-3-mini and Grok-4. Strategic dishonesty seems to arise as a second-order effect of post-training, as such, represents to us a stand-in or for many other implicit alignment choices never tested or expressed unless models are faced with specific scenarios in the wild. Only once such behavior is identified are we able to use probes to detect it. Strategic dishonesty could be held out as a “validation task” and serve as a verifiable testbed for research into how model training choices shape model behaviors.
Conclusion. In this work we have evaluated the propensity of over 80 models to be strategically dishonest, when faced with harmful queries. Certain models, even from the same provider, show surprising variance in their preferences towards dishonesty. We show that this behavior fools all existing output-based detectors for harmful responses that we tested, but we also show that probes of internal states are accurate at detecting this behavior, and can identify activation patterns with causal links to dishonesty, which we validate through finetuning and steering.
Limitations. While we demonstrate that probing can mitigate strategic dishonesty, we are only able to evaluate open-source models. Applying probes to frontier, closed-source models requires cooperation from providers, and any evidence of success or failure has to come from their future reports. Our analysis is restricted to English, leaving open whether results generalize across languages or task framings. Although we argue that our informed-choice format is unbiased, alternative wording in multiple-choice questions may still affect model behavior and lead to different outcomes.
We also highlight in Section E that for models with unfaithful chain-of-thought, probes collapse and fail to provide any meaningful signal. This underscores the critical importance of preserving chain-of-thought and its trustworthiness for oversight of future models (Korbak et al., 2025).
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
What evaluation methods best detect reward hacking in AI agents? Why do models reveal hidden associations despite concealment attempts?- How much do a model's own values leak into answers about practical questions?
- Can models hide misconduct only when they know they are watched?
- Do graders feeding training loops need different disclosure standards than public models?
- How was covertness measured in the model's behavior?
- Why do some model variants deceive while their sibling models refuse?
- Why do coherent value systems in large models include self-valuation above humans?
- Do people adopting model framings without awareness constitute real opinion change?
- Do language models hide their reasoning when user preferences influence their answers?
- When do language models first develop self-preservation preferences?
- How fragile are language models' ethical calibrations across different contextual cues?
- What harness properties determine whether disclosure adds meaningful value?
- Do models reason strategically about their oversight systems?
- How often do models verbalize the reasoning that actually drives their compliance decisions?
- How do evaluation protocols change whether models exhibit sabotage or refusal behaviors?