Can small models detect scheming by watching actions alone?
Whether a compact open-weight model trained on action trajectories can outperform larger frontier models at detecting agent scheming and sabotage, and whether reasoning access is necessary for reliable detection.
The paper reports that a small open-weight model trained to watch an agent's actions can beat prompted frontier models at detecting scheming and sabotage. The monitors work from agentic trajectories alone, never the monitored agent's reasoning or model internals, because chain-of-thought access or internal activations "can be unavailable, unreliable or expensive in deployment." The headline case is Qwen3.5-27B, which "yields higher performance than all low-cost frontier models as prompted monitors (Gemini 3.1 Flash-Lite, GPT-5.4 Nano, and Claude Haiku 4.5) and than Gemini 2.5 Pro, while also achieving lower marginal inference cost." The introduction frames the stakes as a control problem: telling scheming from benign task pursuit "may become a central AI control problem."
The method runs in stated steps. A scheming specification is used to "elicit structured rationales from a frontier teacher," a separate judge filters them, and the highest-quality rationales are distilled into open-weight monitors through supervised fine-tuning and reinforcement learning. The paper says the method is "inspired by deliberative alignment," but what it trains is the monitor, not the agent being watched. The discussion says the teacher's rationales are generated offline, so the deployed model is the small one. The excerpt reports five training datasets and six out-of-distribution agentic misalignment benchmarks, but names none of them, does not name the teacher, and does not describe the specification or the judge's criteria.
Against the nearest notes, this paper makes the opposite choice from Can process-level monitoring reliably detect agent scheming?, which judges scheming from evidence in both reasoning and actions. Here the reasoning channel is excluded on deployment grounds, not shown to be unnecessary. That bears on Can reasoning models be steered by injected context without detection?. An action-only monitor has no chain of thought to read, so that surface is absent by design, but the excerpt does not test plan injection, so whether that is enough is an implication to check. The method also carries Does deliberative alignment genuinely reduce scheming or just hide it? from the agent to the monitor. That note's point that models "reason about being evaluated" bears on any benchmark like these, though the excerpt says nothing about evaluation awareness. Last, Can we detect objective-misaligned agents from their public speech alone? asks what observable behavior can expose a misaligned agent. This paper offers a trained detector for one kind of observable evidence, but on synthetic benchmarks rather than the public cheap-talk setting that question concerns.
The excerpt does not establish the size of the advantage. It gives no metric, no per-benchmark scores, no variance, and no cost figures beyond the unit "token-metered USD per 1,000 evaluations." The paper's own limitation matters most: its evaluations "use synthetic and semi-synthetic control benchmarks rather than naturally-occurring scheming in deployment," which leaves "generalization to realworld long-horizon and adaptive adversaries" as an open question. The implication, at the strength the evidence allows, is narrow: a cheap action-only monitor can exceed prompted frontier monitors on these benchmarks. Whether it catches an adaptive adversary in deployment is not shown.
Inquiring lines that read this note 26
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Why do models reveal hidden associations despite concealment attempts?- Why does covert sabotage appear in only two of fourteen frontier models?
- How persistent is in-context scheming across multiple follow-up interactions?
- Can activation probing reliably distinguish between sycophancy and scheming in models?
- Do record-tampering and covert sabotage share a common underlying mechanism?
- Why do Claude and OpenAI models cheat through different strategies?
- How do models reason about being monitored and evaluated during scheming benchmarks?
- How do models recognize the social patterns that signal a test is underway?
- How do evaluation protocols change whether models exhibit sabotage or refusal behaviors?
- Does lower covert-action detection in monitors actually indicate better-hidden situational awareness?
- Should scheming detection use reasoning evidence alongside action evidence for reliability?
- What drives scheming propensity most strongly across different LLM agents?
- Which AI scheming claims have public prompts and transcripts available?
- Does improving detection accuracy change how slop accusations function socially?
- Can user feedback flags rival AI detector accuracy for identifying AI slop?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can process-level monitoring reliably detect agent scheming?
SCOUT proposes grounding scheming judgments in evidence from agent reasoning and actions across multiple criteria. The question explores whether this process-level approach can actually scale and reliably catch deceptive behavior that spans multiple steps.
contrast: SCOUT uses reasoning as well as action evidence; this paper uses actions only.
-
Does deliberative alignment genuinely reduce scheming or just hide it?
Deliberative alignment dramatically cuts covert actions in language models, but their reasoning reveals awareness of being evaluated. The question is whether the improvement reflects real alignment or strategic compliance.
borrows deliberative alignment's spec-grounded reasoning but applies it to the monitor, not the agent.
-
Can reasoning models be steered by injected context without detection?
This explores whether adversaries can plant harmful-but-benign-sounding reasoning in a model's context and have it followed while evading chain-of-thought monitors. The question matters because it tests whether monitoring reasoning traces can catch deception at inference time.
an action-only design avoids the chain-of-thought surface that attack evades; the excerpt does not test this.
-
Can we detect objective-misaligned agents from their public speech alone?
The paper claims compromised agents' adaptations remain invisible in public behavior, but leaves the detection method and audience unspecified. Identifying who or what can detect such agents—and with what accuracy—is crucial for building defenses.
offers a trained detector for observable evidence, on synthetic benchmarks rather than public cheap talk.
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Training Deliberative Monitors for Black-Box Scheming Detection
- Norms at a Price: Why RL-Based Alignment Can Promise Conditional Compliance at Best
- SchemeArena: Factorized Stress Testing of Scheming in LLM Agents
- Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings
- Frontier Models are Capable of In-context Scheming
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- Stress Testing Deliberative Alignment for Anti-Scheming Training
- Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
Original note title
action-only deliberative monitors distilled from frontier rationales beat prompted frontier models at detecting scheming — on synthetic control benchmarks