Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings
Chain-of-thought (CoT) monitoring is increasingly treated as an important safety layer for frontier reasoning models. Most monitorability evaluations study explicitinfluence settings: setups where the prompt directly incentivizes the model to hide something, e.g., by instructing it to perform a hidden side-task. A complementary axis for CoT-monitor evaluations is implicit-influence settings, where the prompt contains no instruction to hide, but the model’s behavior is still shaped by features of the task or context, e.g. an irrelevant detail about a candidate that biases a hiring rating. We introduce the first benchmark that directly compares CoT monitorability under the two regimes. We test how model behavior changes in the presence of a nudge to choose a particular options. The nudge is delivered either as a casual aside (implicit), or as a direct instruction to act on the nudge and to conceal having done so (explicit). The benchmark spans four task formats (binary choice, numeric rating, multiple-choice QA, open-ended coding) and seven frontier extended-thinking models. Under explicit influence, a CoT monitor detects 60–94% of behavior shifts: even models instructed to conceal it leak the instruction into their CoT.
Introduction. Chain-of-thought (CoT) monitoring, using a model’s reasoning trace to detect problematic behavior, is increasingly treated as an important safety layer for frontier AI. A growing body of work supports this optimism. Models attempting covert side-tasks in SHADE-Arena frequently reveal their intent in their reasoning traces [Zolkowski et al., 2025], reasoning models struggle to control the content of their own CoT [Chen Yueh-Han et al., 2026], and across a range of sabotage and reward-hacking benchmarks, frontier reasoning models are found to be fairly monitorable [Guan et al., 2025, Baker et al., 2025]. Existing work covers a wide variety of environments, with different monitorability properties. We propose classifying monitorability evaluations by how the behavior that the monitor must detect is created.
Discussion / Conclusion. Implicit influence is a harder regime for CoT monitoring. Across settings, the matched implicitexplicit comparison shows detection 41–46 percentage points lower under implicit delivery in two settings and 11 points lower in a third, with Setting 4 the exception only because its implicit condition is already at ceiling. In particular, the hard regime is not the adversarial one: the explicit variants instruct the model to conceal its reason, yet models leak their compliance into the CoT, consistent with prior findings that models are poor at hiding intent on demand [Chen Yueh-Han et al., 2026]. This suggests that explicit-influence settings may give a poor estimate of CoT monitorability in the wild. Making reasoning load-bearing helps, but doesn’t close the gap. Making the reasoning loadbearing for the nudge (acting on it requires tracing a multi-hop chain the model cannot compute in a single forward pass) recovers much of the lost detection (12%→51% on binary preference), consistent with Emmons et al. [2025]. However this comes with two caveats.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Why does polished AI output gain credibility despite fundamental verifiability problems? How does awareness of evaluation context influence model behavior? Can reasoning traces reveal actual model reasoning versus plausible output?- Can corrupted reasoning traces be reliably distinguished from correct ones?
- How much accuracy is preserved when removing explanatory layers from reasoning traces?
- Can reasoning traces prove models are actually reasoning versus mimicking?
- Does anonymizing reasoning traces harm the quality of model outputs?
- Why do we measure reasoning quality by reading visible chains?
- How do adversarial triggers bypass the protections of longer reasoning chains?
- How can simple prompt injection attacks extract reasoning trace content?
- Can synthesized explanations be more auditable than winning-chain explanations?
- What infrastructure could replace search for verifying AI outputs?