SYNTHESIS NOTE
Topics›Agents Multi Architecture›this note

Can inspecting generated workflows catch planning-time attacks?

Does examining a workflow after it's created catch attacks that corrupt the planning signals upstream? This matters because if contamination enters earlier, downstream inspection might miss malicious intent laundered into legitimate-looking structure.

Synthesis note · 2026-05-28 · sourced from Agents Multi Architecture

A defense can only catch what it can see, and where it looks determines what it can catch. Because FLOWSTEER biases the planning signals from which the workflow is generated, any defense that inspects only the resulting workflow examines an artifact that is already compromised. The malicious intent has been laundered through the planner into legitimate-looking roles, dependencies, and routing — by the time the workflow exists, the contamination is no longer visibly malicious. This is why the paper introduces FLOWGUARD as an input-side defense: it strengthens the planning boundary by separating task, methodological, and framing intents, then reframes workflow-contaminating cues while preserving the original task objective, reducing malicious success by up to 34 percent without degrading prompt utility.

The general principle is about defense placement, not defense strength. Moving inspection upstream — to the point where intent is parsed but before organization is committed — catches a class of attack that downstream inspection structurally cannot. The counterpoint is that input-side defense risks false positives that suppress legitimate methodological guidance, which is exactly why FLOWGUARD separates intent types rather than filtering wholesale. This matters because it reframes MAS security as a question of where the trust boundary sits: the safest place to intervene is the boundary between instruction and organization, not the organization itself.

Inquiring lines that read this note 21

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do standardized protocols improve multi-agent coordination and reliability? How does misalignment propagate through agent communication networks? Can single-point security defenses protect multi-agent systems from multi-step attacks? What attack surfaces do reasoning traces and chains introduce? Can local safety checks guarantee system-level behavioral safety? Why do locally safe actions create system-level safety gaps? How do we enforce security boundaries in evaluation environments? What determines whether deployed AI systems can actually be stopped in practice? Can reasoning traces and behavior monitoring reliably detect hidden AI scheming? How do evaluation practices shape which failures stay visible? How can we detect and prevent harm propagation through multi-agent delegation workflows?

Related concepts in this collection 8

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
19 direct connections · 119 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

defenses that inspect only the generated workflow miss attacks that bias the upstream planning signal