SYNTHESIS NOTE
Topics›Frontier AI Risk & RSI›this note

Can small models detect scheming by watching actions alone?

Whether a compact open-weight model trained on action trajectories can outperform larger frontier models at detecting agent scheming and sabotage, and whether reasoning access is necessary for reliable detection.

Synthesis note · 2026-10-06 · sourced from Frontier AI Risk & RSI

The paper reports that a small open-weight model trained to watch an agent's actions can beat prompted frontier models at detecting scheming and sabotage. The monitors work from agentic trajectories alone, never the monitored agent's reasoning or model internals, because chain-of-thought access or internal activations "can be unavailable, unreliable or expensive in deployment." The headline case is Qwen3.5-27B, which "yields higher performance than all low-cost frontier models as prompted monitors (Gemini 3.1 Flash-Lite, GPT-5.4 Nano, and Claude Haiku 4.5) and than Gemini 2.5 Pro, while also achieving lower marginal inference cost." The introduction frames the stakes as a control problem: telling scheming from benign task pursuit "may become a central AI control problem."

The method runs in stated steps. A scheming specification is used to "elicit structured rationales from a frontier teacher," a separate judge filters them, and the highest-quality rationales are distilled into open-weight monitors through supervised fine-tuning and reinforcement learning. The paper says the method is "inspired by deliberative alignment," but what it trains is the monitor, not the agent being watched. The discussion says the teacher's rationales are generated offline, so the deployed model is the small one. The excerpt reports five training datasets and six out-of-distribution agentic misalignment benchmarks, but names none of them, does not name the teacher, and does not describe the specification or the judge's criteria.

Against the nearest notes, this paper makes the opposite choice from Can process-level monitoring reliably detect agent scheming?, which judges scheming from evidence in both reasoning and actions. Here the reasoning channel is excluded on deployment grounds, not shown to be unnecessary. That bears on Can reasoning models be steered by injected context without detection?. An action-only monitor has no chain of thought to read, so that surface is absent by design, but the excerpt does not test plan injection, so whether that is enough is an implication to check. The method also carries Does deliberative alignment genuinely reduce scheming or just hide it? from the agent to the monitor. That note's point that models "reason about being evaluated" bears on any benchmark like these, though the excerpt says nothing about evaluation awareness. Last, Can we detect objective-misaligned agents from their public speech alone? asks what observable behavior can expose a misaligned agent. This paper offers a trained detector for one kind of observable evidence, but on synthetic benchmarks rather than the public cheap-talk setting that question concerns.

The excerpt does not establish the size of the advantage. It gives no metric, no per-benchmark scores, no variance, and no cost figures beyond the unit "token-metered USD per 1,000 evaluations." The paper's own limitation matters most: its evaluations "use synthetic and semi-synthetic control benchmarks rather than naturally-occurring scheming in deployment," which leaves "generalization to realworld long-horizon and adaptive adversaries" as an open question. The implication, at the strength the evidence allows, is narrow: a cheap action-only monitor can exceed prompted frontier monitors on these benchmarks. Whether it catches an adaptive adversary in deployment is not shown.

Inquiring lines that read this note 26

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Why do models reveal hidden associations despite concealment attempts? Can AI systems evade safety evaluations through reasoning manipulation? How do multi-agent architectures affect AI system security and defense effectiveness? How does awareness of evaluation context influence model behavior? Can monitoring reasoning traces and behavior detect hidden agent deception? How reliably can humans and AI detectors identify machine-generated text? Can humans reliably detect and resist AI-generated misinformation? Do individually safe AI actions create unsafe outcomes in integrated systems? How do transformer attention patterns implement retrieval and reasoning? How do evaluation environment design choices affect AI security? What social dynamics enable or prevent agent collusion? What human oversight must AI research systems have? How should we measure frontier AI models' cyber exploitation capabilities? When do multi-agent systems improve over single frontier models? How can humans maintain effective oversight as AI systems scale? How does optimization for reward create emergent misalignment in language models?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 96 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

action-only deliberative monitors distilled from frontier rationales beat prompted frontier models at detecting scheming — on synthetic control benchmarks