When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors
While chain-of-thought (CoT) monitoring is an appealing AI safety defense, recent work on “unfaithfulness” has cast doubt on its reliability. These findings highlight an important failure mode, particularly when CoT acts as a post-hoc rationalization in applications like auditing for bias. However, for the distinct problem of runtime monitoring to prevent severe harm, we argue the key property is not faithfulness but monitorability. To this end, we introduce a conceptual framework distinguishing CoT-as-rationalization from CoT-as-computation. We expect that certain classes of severe harm will require complex, multi-step reasoning that necessitates CoT-as-computation. Replicating the experimental setups of prior work, we increase the difficulty of the bad behavior to enforce this necessity condition; this forces the model to expose its reasoning, making it monitorable. We then present methodology guidelines to stress-test CoT monitoring against deliberate evasion. Applying these guidelines, we find that models can learn to obscure their intentions, but only when given significant help, such as detailed human-written strategies or iterative optimization against the monitor.
Introduction. As today’s language models reason in natural language, chain-of-thought (CoT) monitoring is an intuitive safety defense. The idea is simple: if you want to prevent bad behavior, then look for bad reasoning in the model’s CoT. Despite this intuitive appeal, the reliability of CoT monitoring has been cast into doubt by a significant body of recent research on “unfaithfulness” (Arcuschin et al., 2025; Chen et al., 2025; Chua and Evans, 2025; Lanham et al., 2023; Turpin et al., 2023). This work shows that models can fabricate justifications, omitting the true causes of their decisions. This is a critical issue for many applications; for instance, a fabricated CoT could undermine an audit that’s looking for bias in a model used for hiring decisions (Karvonen and Marks, 2025). However, for the distinct problem of runtime monitoring for severe harm, we argue that unfaithful rationalizations are not a dealbreaker. The ultimate property we need for safety is monitorability: whether we can use the CoT to determine if a model is taking dangerous actions (Baker et al., 2025).
Discussion / Conclusion. Our key reason for optimism about CoT monitorability is that CoT might be necessary to cause severe harm, as discussed in Section 2. Compared to the impression that previous work on unfaithfulness may generate, we think this leads to an overall more positive prospect for CoT monitoring, as discussed in Section 3. However, this relies to a large extent on conceptual arguments about whether necessity will hold. Our stress-testing results in Sections 5 to 7 address a separate concern: if CoT is necessary for a side task, could models still evade a CoT monitor if they wanted to? We show that current models can at best evade monitors with significant help in our environments. We think they likely would not be able to evade monitors under realistic deployment conditions. On the other hand, we do find signs of life for monitor evasion in Sections 5 and 6. We thus can’t draw confident conclusions about if or when future models might be able to evade CoT monitors.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Does chain-of-thought reasoning reveal how models actually think or merely imitate reasoning? How do individually-safe actions create collectively-unsafe outcomes?- What false-positive rates do chain-of-thought safety monitors achieve?
- How do current safety benchmarks miss pragmatic alignment failures?
- How does Goodhart's Law apply when safety measures become optimization targets?
- How strong is the claim that severe harms require multi-step reasoning?
- Can models deliberately obfuscate reasoning to defeat chain-of-thought safety monitors?
- Can safety evaluations miss behavioral effects by only measuring semantic shifts?
- What safety protections work when simulators have access to real APIs?
- Which AI safety problems lack the scalar metrics autoresearch requires?
- What calibration corrections can reduce LLM judge bias in automated evaluation pipelines?
- What happens when LLMs grade other LLMs in closed evaluation loops?
- What design principles prevent error cascades in multi-step evaluation systems?
- How do autonomous pipelines identify and fix silent bugs in data pipelines?
- What specific failure modes must evaluation catch before deploying action-capable systems?