Can inspecting generated workflows catch planning-time attacks?
Does examining a workflow after it's created catch attacks that corrupt the planning signals upstream? This matters because if contamination enters earlier, downstream inspection might miss malicious intent laundered into legitimate-looking structure.
A defense can only catch what it can see, and where it looks determines what it can catch. Because FLOWSTEER biases the planning signals from which the workflow is generated, any defense that inspects only the resulting workflow examines an artifact that is already compromised. The malicious intent has been laundered through the planner into legitimate-looking roles, dependencies, and routing — by the time the workflow exists, the contamination is no longer visibly malicious. This is why the paper introduces FLOWGUARD as an input-side defense: it strengthens the planning boundary by separating task, methodological, and framing intents, then reframes workflow-contaminating cues while preserving the original task objective, reducing malicious success by up to 34 percent without degrading prompt utility.
The general principle is about defense placement, not defense strength. Moving inspection upstream — to the point where intent is parsed but before organization is committed — catches a class of attack that downstream inspection structurally cannot. The counterpoint is that input-side defense risks false positives that suppress legitimate methodological guidance, which is exactly why FLOWGUARD separates intent types rather than filtering wholesale. This matters because it reframes MAS security as a question of where the trust boundary sits: the safest place to intervene is the boundary between instruction and organization, not the organization itself.
Inquiring lines that read this note 21
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do standardized protocols improve multi-agent coordination and reliability? How does misalignment propagate through agent communication networks?- Why does workflow position amplify malicious signals downstream?
- What happens when planning signals get contaminated before reaching a downstream agent?
- How does workflow position amplify or suppress malicious signals?
- How does position in a workflow amplify or suppress harmful agent behavior?
- What makes planning-time attacks structurally invisible to downstream inspection?
- How do workflow-inspecting defenses fail when contamination enters at planning time?
- How do defenses that inspect planning signals compare to workflow-level validation?
- Do per-hop inspection gates miss attacks that bias upstream planning signals?
- Where do workflow inspection defenses fail against upstream planning attacks?
- Can human inspection of auto-generated workflows catch harmful or incorrect API compositions?
- Can a system pass all local checks while the overall workflow still fails?
- Why do workflow-level defenses catch attacks that single-skill inspection cannot detect?
- Why can every step pass its local check while a workflow still fails?
Related concepts in this collection 8
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can we defend RAG systems from corpus poisoning without retraining?
Explores whether retrieval-time defenses can catch and block poisoned documents before they reach the generator, without expensive retraining cycles. Matters because corpus updates outpace model retraining in production RAG systems.
parallel principle that the right defense sits upstream of where the harm becomes visible
-
How do adversarial traps target different layers of AI agents?
As AI agents browse the web, attackers can exploit their perception, reasoning, memory, actions, and coordination in distinct ways. Understanding these attack vectors is crucial for building robust agent defenses.
locating defenses depends on which trap category an attack belongs to
-
Can prompts alone reshape multi-agent workflows without system access?
Explores whether attackers can compromise planner-executor multi-agent systems by manipulating the planning prompt itself, without touching agents, tools, or infrastructure. Matters because it identifies a previously overlooked attack surface that existing defenses don't address.
same FLOWSTEER work; names the planning-time attack surface that this note argues downstream workflow inspection structurally cannot see
-
How does a signal's position in a workflow change its influence?
Multi-agent systems may amplify or suppress malicious signals based on where they enter the workflow. Understanding position-dependent propagation could reveal which nodes are most critical to defend.
explains the propagation mechanism that makes upstream contamination look legitimate by the time it reaches the workflow this note says is inspected too late
-
Can chain-of-thought monitors detect reasoning that originates elsewhere?
When language models work inside pipelines that inject reasoning from retrieved documents, planners, or other agents, monitoring systems may evaluate paraphrased external reasoning as if it were the model's own thinking. This raises questions about what monitors can actually detect.
the same placement problem at a CoT monitor: it reads the downstream artifact, the trace, while the contamination entered upstream in context and was paraphrased into clean-looking words
-
Do internal agent hops in pipelines need security monitoring?
Multi-agent systems route data between planner, worker, verifier, and synthesizer components. Current defenses only guard user input at the entry point, leaving inter-agent channels unmonitored—but is this a real vulnerability or does downstream safety suffice?
adds how many channels a defense covers to this note's argument about where it looks; a filed tension in ops/tensions/ sets one upstream boundary against a gate on every hop
-
Should sanitizers re-score their compressed output before passing it?
When a gate compresses text to remove harmful content, does it need to verify the compressed remainder is still safe? The question matters because ChannelGuard's current approach assumes position—that payloads are removed—without checking if what remains passes the same safety test.
FLOWGUARD reframes cues before the planner sees them, a rewrite-before-pass step; whether the rewritten output is inspected again is the question that note asks of a compress gate
-
Which attack and defense numbers came from filtered backends?
Research figures on multi-agent attacks and defenses may have been measured behind provider filters that silently shaped the results. Understanding which numbers had this filter dependency is critical for interpreting their real-world strength.
the 34 percent reduction is a defense figure the audit would label with its backend and filter status
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- FLOWSTEER: Prompt-Only Workflow Steering Exposes Planning-Time Vulnerabilities in Multi-Agent LLM Systems
- Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection
- SafeFlow: Semantic Information-Flow Control for Blocking Malicious Propagation in Multi-Agent Systems
- ChannelGuard: Safe Models Do Not Compose into Safe Multi-Agent Systems
- A Black Box for Agentic Processes: Blockchain-Anchored Evidence for AI Agent Communication, Human Oversight, and GRC Audits
- PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?
- SchemeArena: Factorized Stress Testing of Scheming in LLM Agents
- COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation
Original note title
defenses that inspect only the generated workflow miss attacks that bias the upstream planning signal