INQUIRING LINE

Can an unseen safety filter on an AI company's servers quietly make published attack results look better or worse than the model really is?

Can server-side filters silently provide safety credit to published attack results?

This explores whether safety filters that AI providers run on their own servers, out of the researcher's view, can quietly block some attacks during an experiment, so that a paper's attack-success or defense numbers look better or worse than the model itself really is.


This explores whether a provider's hidden server-side filter can tilt a published security result: an attack looks weaker, or a defense looks stronger, because an unseen filter did part of the blocking. The collection's most direct answer is a warning about how results are reported, not a measurement. One audit of multi-agent attack and defense papers found that most published attack-success and defense-gain percentages never say whether they were measured behind a server-side filter Which attack and defense numbers came from filtered backends?. When that label is missing, a reader can't tell whether a figure like '40% of attacks succeeded' describes the model, the paper's proposed defense, or a provider filter that nobody mentioned. So yes, the safety credit can go to the wrong thing. The collection flags this as a gap but does not yet measure how large the distortion is.

The credit can also go to a different model entirely. The layer that decides which model handles a request is itself an attack surface, and its provenance is often unverified. Requests can end up at a different model than intended, and safety checks can run against the wrong identity Can attackers manipulate which model handles a request?. A researcher who calls an API endpoint may assume they are testing one model with one set of protections, while the provider's infrastructure quietly changes both. A filter is only one of several hidden layers that sit between 'I ran the attack' and 'here is what happened.'

A related argument concerns what filters can actually catch. A filter judges a single output at a single moment, but an agent's risk is spread across its memory, the documents it retrieves, its tool calls, and what it can reach in its environment Can a model-level filter truly contain an agent with environment access?. Skill scanners show the same weakness. They score each piece separately, so an attacker can make each piece look harmless while the attack chain as a whole still works, reaching 96% success across six scanners Can attackers evade skill scanners by refining individual skills?. Put these together and a filtered result can mislead twice. It can understate how well an attack does against the bare model, and it can overstate protection in agent deployments where the filter only sees fragments.

The most useful idea comes from a neighboring field: cheating on benchmarks. Researchers there faced a similar problem, where a score alone could not show what actually happened during a run. Existing defenses against reward hacking produce no reusable evidence that a specific run stayed within the rules Do current reward-hacking defenses provide reusable evidence of safety?. The proposed fix is to record infrastructure evidence alongside the score, so that a claim of valid completion rests on a record of what happened Can infrastructure evidence replace terminal scores in benchmark validation?. Logging the points in a run where an agent gains or uses authority can also separate tasks where a hack was merely possible from runs where it was actually used Can runtime instrumentation distinguish hacking exposure from actual exploitation?. Applied to security papers, the same approach would mean reporting which filters were active, which model actually answered, and which attempts were blocked by infrastructure instead of the model. Each result would then carry its own record of how it was produced.

The gap is worth noting directly. The collection names the problem and offers a tested way to think about fixing it. It does not contain a study that runs the same attacks with filters on and off and measures the difference. Until such a study exists, the safest reading of any unlabeled attack or defense number from a commercial API is that it measures the whole system the provider runs, not the model alone.


Sources 7 notes

Which attack and defense numbers came from filtered backends?

Attack success and defense gain percentages across the vault are reported without disclosing whether they were measured behind server-side filters. This omission makes it impossible to determine whether results reflect true model behavior or filtered outcomes.

Can attackers manipulate which model handles a request?

The layer deciding which model handles a request sits beneath prompt-level defenses and is vulnerable to manipulation and unverified provenance. Attackers can exploit routing to send requests to weaker models or cause safety measures to operate on the wrong identity.

Can a model-level filter truly contain an agent with environment access?

A filter judges a single output at one point in time; an agent's risk spreads across memory, retrieved content, tool calls, and environmental reach. Containment requires controlling what an agent can touch, not just what it says now.

Can attackers evade skill scanners by refining individual skills?

ColluSkill combines chain planning with scanner-feedback refinement to reach 96% average attack success. The approach works because scanners score skills individually, allowing feedback to reduce suspicion per skill while chain-level semantics remain intact.

Do current reward-hacking defenses provide reusable evidence of safety?

Existing defenses rely on task-specific patches, prompt instructions, or post-hoc detectors, but none provide reusable evidence that a concrete run remained within its evaluation boundary. Even working defenses do not give operators a portable record of integrity.

Show all 7 sources
Can infrastructure evidence replace terminal scores in benchmark validation?

BenchShield enables benchmark operators to issue claims about valid task completion grounded in recorded infrastructure evidence rather than terminal scores alone. This shifts from a single number to a verifiable claim about whether an agent followed the intended evaluation path.

Can runtime instrumentation distinguish hacking exposure from actual exploitation?

Infrastructure-side recording of authority-bearing transitions distinguishes tasks that merely expose a hacking vector from runs that actually exercise one. This separation prevents every score from an exposed task being automatically suspect.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.