INQUIRING LINE

Can one tidy 'how safe is our monitoring' score hide the specific spots where someone slipped past it?

Can aggregate monitorability metrics miss targeted evasion in specific domains?

This explores whether a single overall score for how well we can monitor an AI system ("the monitor catches X% of bad behavior") can hide blind spots where a model, or an attacker, evades detection in one particular area.


This explores whether one headline number for how monitorable an AI system is can hide pockets where evasion succeeds. The corpus has no study that breaks a published monitorability score down by domain to find hidden gaps. What it does have is a recurring pattern from several angles: monitors that look fine on average often fail when someone works around the way they slice up behavior. Most of the blind spots come from what the monitor treats as one unit of review, more than from which subject area is involved.

The clearest example is Can attackers evade skill scanners by refining individual skills?. Six different skill scanners each score agent skills one at a time. An attacker uses the scanners' own feedback to make each skill look harmless while the harmful plan lives in how the skills chain together, and reaches 96% attack success. Each scanner can perform well on its own benchmark and still miss this completely, because the harm isn't in any single item it scores. Can task decomposition hide harmful intent across agents? finds the same thing in multi-agent systems. Splitting a task into roles, which is what makes these systems useful, also lets a harmful goal be divided into steps that each look benign. The harm only shows up when the pieces are combined. Should defence units span multiple executions and agents? draws the practical conclusion: monitor "coordination episodes," meaning groups of linked actions that get revised as evidence comes in, rather than isolated actions. Does added monitoring improve protection at acceptable cost? proposes a test of whether that kind of grouping actually helps at equal review cost. Notably, it reports no results yet.

A quieter warning comes from difference-of-means-vectors-are-similarly-effective-to-llm-monitors-but-virtually. A cheap detector that reads the model's internal activations roughly matches an expensive LLM monitor overall. But broken down by model, it catches 3.1% more reward hacks on one model and 7.9% fewer on another. "Similarly effective" is a true summary that hides real differences underneath, and that's the sort of variation a single aggregate score smooths over. Can distance alone rank which substrates resist reward hacking? puts this in formal terms. How exposed a system is depends on where the evaluator's errors sit relative to the behaviors the system can actually reach. As that location shifts, the ranking shifts too, so no general vulnerability ranking holds up across settings.

A second family of evasion goes around the monitor rather than through it. Can attackers manipulate which model handles a request? shows that attackers can steer requests to weaker models, or make safety checks run against the wrong identity. A monitor's measured accuracy then says nothing about the traffic it never sees. The BenchShield work points to one way out: stop trusting a single final score and record what actually happened in the infrastructure. Can runtime instrumentation distinguish hacking exposure from actual exploitation? uses those records to tell tasks that merely made a hack possible apart from runs that actually used one. Can infrastructure evidence replace terminal scores in benchmark validation? turns this into verifiable claims about how a task was completed, not just how well. Can scoped agents reliably judge semantic hacks in runtime analysis? adds tightly scoped auditors that must cite evidence for their judgments, though how reliable those auditors are hasn't been measured yet.

The main takeaway: an aggregate score can hide blind spots in particular subject areas, but the bigger risk is that it averages over the wrong unit. Evasion that works by splitting harmful work into innocent-looking pieces, or by rerouting requests past the monitor, never appears as low accuracy on individual items. So when you see a monitorability number, ask what unit was scored (single actions, whole episodes, or a single model's traffic) and whether anyone looked at how results vary underneath the average.


Sources 10 notes

Can attackers evade skill scanners by refining individual skills?

ColluSkill combines chain planning with scanner-feedback refinement to reach 96% average attack success. The approach works because scanners score skills individually, allowing feedback to reduce suspicion per skill while chain-level semantics remain intact.

Can task decomposition hide harmful intent across agents?

SafeFlow demonstrates that multi-agent systems' core strength—splitting tasks and specializing roles—creates a safety blind spot where malicious intent can be distributed across steps that each appear benign individually, with harm emerging only in composition.

Should defence units span multiple executions and agents?

The operational unit of defence should be a set of actions linked by observed transfers, task authority, and response history, with membership revised as evidence accumulates. Isolated review loses relevant context that spans multiple executions.

Does added monitoring improve protection at acceptable cost?

The paper designs a controlled comparison of isolated actions, rolling windows, known groups, and prospectively discovered episodes at equal review cost and false-alert workload, but the excerpt provides no empirical results showing whether added monitoring improves protection.

Can distance alone rank which substrates resist reward hacking?

A distance-dependent error bound and capacity ordering are formal statements about limits, not predictions of what systems find. Where evaluator errors sit among reachable behaviors and search effectiveness determine actual exposure, which shifts as the scoring defect location changes.

Show all 9 sources
Can attackers manipulate which model handles a request?

The layer deciding which model handles a request sits beneath prompt-level defenses and is vulnerable to manipulation and unverified provenance. Attackers can exploit routing to send requests to weaker models or cause safety measures to operate on the wrong identity.

Can runtime instrumentation distinguish hacking exposure from actual exploitation?

Infrastructure-side recording of authority-bearing transitions distinguishes tasks that merely expose a hacking vector from runs that actually exercise one. This separation prevents every score from an exposed task being automatically suspect.

Can infrastructure evidence replace terminal scores in benchmark validation?

BenchShield enables benchmark operators to issue claims about valid task completion grounded in recorded infrastructure evidence rather than terminal scores alone. This shifts from a single number to a verifiable claim about whether an agent followed the intended evaluation path.

Can scoped agents reliably judge semantic hacks in runtime analysis?

BenchShield constrains audit agents by limiting their remit, fixing the artifacts they see, and requiring evidence citation. This positions infrastructure records as unchallengeable checks and audit judgments as the arguable step after them, though reported reliability remains unquantified.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.