Can prompts alone reshape multi-agent workflows without system access?
Explores whether attackers can compromise planner-executor multi-agent systems by manipulating the planning prompt itself, without touching agents, tools, or infrastructure. Matters because it identifies a previously overlooked attack surface that existing defenses don't address.
The flexibility that makes planner-executor multi-agent systems attractive is also their weakness. When a planner converts a prompt into subtasks, roles, dependencies, and routing paths, the prompt is not merely a request — it is the blueprint from which the entire collaboration is constructed. FLOWSTEER demonstrates that an attacker who never touches agents, tools, memory, or inter-agent messages can still steer behavior, because the planning step happens before any of that infrastructure is invoked. A single crafted prompt can bias how the workflow forms in the first place, raising malicious success by up to 55 percent over naive prompting and transferring across MAS setups even under black-box topology inference.
This reframes where multi-agent safety lives. Most existing defenses inspect the artifacts of coordination — the generated workflow, the messages exchanged, the tool calls made. But if the contamination enters at workflow formation, those defenses arrive too late. The attack surface is not the running system; it is the organizational act of deciding who does what and in what order. The counterpoint is that this requires the planner to be promptable at all — fully fixed pipelines are immune — but fixed pipelines forfeit the adaptive coordination that motivates planner-executor designs. This matters because it identifies workflow formation as a distinct security frontier, one that grows more exposed precisely as multi-agent systems become more flexible.
Inquiring lines that read this note 103
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can single-point security defenses protect multi-agent systems from multi-step attacks?- Can message-layer defenses stop prompt injection across multi-agent networks?
- Why does attack generation scale faster than defense engineering?
- What makes planning-time attacks structurally invisible to downstream inspection?
- How do workflow-inspecting defenses fail when contamination enters at planning time?
- Can fixed pipelines eliminate planning-time attacks by sacrificing adaptive coordination?
- Can existing web security defenses protect agents from content manipulation?
- Does prompt hardening equally protect single and multi-agent web systems?
- What makes the Telephone Loop attack specific to agent delegation?
- How do defenses that inspect planning signals compare to workflow-level validation?
- Do per-hop inspection gates miss attacks that bias upstream planning signals?
- How do unmonitored channels between pipeline agents enable security gaps?
- Should defense against coordinated intrusion span multiple execution episodes?
- Does amplifying a single-actor failure require different security defenses than preventing it?
- Can input-boundary defenses guard unmonitored channels between agent hops?
- How do multi-step exploitation chains make agent containment harder to achieve?
- Why do input-boundary defenses fail in planner-worker pipelines?
- Where do workflow inspection defenses fail against upstream planning attacks?
- Can fixed pipelines eliminate planning-time attack surfaces in multi-agent systems?
- Can attackers exploit pooled agent trajectories to identify and bypass defenses?
- What attacks does the agent-specific attack surface decompose into?
- How does prompt hardening work differently in single-agent versus multi-agent systems?
- How much does prompt hardening actually defend multi-agent systems?
- Why do single-agent and multi-agent systems show different defense effectiveness?
- Can prompt hardening reduce signal propagation in multi-agent systems?
- Can defenses at planning boundaries catch attacks that bias upstream instruction signals?
- What are the eight attack vectors used to probe agents in OpenART?
- How do hardened prompts defend against adversarial attacks in multi-agent systems?
- How do manipulative prompts exploit the length-accuracy vulnerability?
- What makes extended chains more vulnerable than standard prompts?
- Why does sandboxed execution matter more than monolithic prompting?
- How can simple prompt injection attacks extract reasoning trace content?
- Do legitimate task signals exploit the same position and framing vulnerabilities as attacks?
- What makes injected plans different from optimization pressure against monitors?
- How can model routing and provenance become an attack surface?
- What makes reasoning-shaped payloads more effective than command-shaped attack prompts?
- Does surface-form query rewriting allow attackers to steer model routing decisions?
- What separates good workflow design from poor workflow design?
- What makes protocols better than free-form prompting for tool coordination?
- Why does agent-to-agent interaction expose identity verification vulnerabilities?
- Can protocol bridges introduce new failure modes or security vulnerabilities?
- Can replanning in multi-agent systems introduce new attack surface or reduce it?
- How does task division in multi-agent design affect security outcomes?
- What attacks are unique to multi-agent systems compared to single agents?
- How do shared artifact stores become security risks in multi-agent systems?
- What makes agent-to-agent messages in multi-agent systems vulnerable to exploitation?
- Where should the trust boundary sit in multi-agent planner systems?
- How does a single compromised agent degrade performance across entire multi-agent pipelines?
- What baseline would prove multi-agent systems are actually less safe?
- What containment risks emerge as agents obtain successive exploit primitives?
- Which agent properties like state retention enable supply-chain and credential vulnerabilities?
- How do agent-to-agent messages bypass defenses on downstream principals?
- Can specialized roles let malicious objectives hide across multiple agents?
- When do multi-agent architectures create more attack surface than single-agent systems?
- Which interaction interfaces do multi-agent systems expose to adversaries?
- How does insider threat differ from external attack in multi-agent systems?
- What vulnerabilities emerge at each hop between agents in a pipeline?
- Can shared memory poisoning compromise multi-agent delegation chains?
- Does quarantining state count as recovery in multi-agent attack scenarios?
- How does protocol mediation affect determinism in agentic function calls?
- Why does pre-computed workflow generation work better than runtime tool discovery for data security?
- Can open agent workflows be modeled as finite event lifecycles?
- What makes a tool schema high-quality enough to prevent agent misuse?
- Why does workflow position amplify malicious signals downstream?
- Why does workflow position amplify malicious signals in multi-agent relay chains?
- How does prompt injection differ from subliminal message propagation in multi-agent networks?
- Do prompt injection attacks propagate behavioral bias across multi-agent networks?
- What happens when planning signals get contaminated before reaching a downstream agent?
- How does workflow position amplify or suppress malicious signals?
- How does workflow position amplify malicious signals in multi-agent systems?
- How does position in a workflow amplify or suppress harmful agent behavior?
- Can human inspection of auto-generated workflows catch harmful or incorrect API compositions?
- Do sequences of individually safe actions collectively violate system-level constraints?
- Can we empirically test whether open models lower barriers to harmful workflows?
- Can individual permissible actions collectively violate system-level constraints?
- Can system prompts alone enforce compliance rules without external enforcement?
- Can delegation prevent silent corruption in long delegated workflows?
- How does shared state convert temporary compromise into persistent inherited risk?
- Can the policy oracle itself be written to by agents in the pipeline?
- Why do workflow-level defenses catch attacks that single-skill inspection cannot detect?
- Why does a single approval point create an easy target for attackers?
- Can individual actions be safe while sequences of them violate system constraints?
- Can attackers assemble harmful outcomes from multiple individually authorized subtasks?
- Can a single authorization policy distinguish licensed delegation from intrusion?
- Can the same tool call be both authorized and unauthorized depending on intent?
- What keeps the task-bound token and policy oracle isolated from poisoning?
- Why do outcome-level metrics fail to reveal contained attacks in multi-agent pipelines?
- What does agent security look like when measured across interaction trajectories?
- What must auditors reconstruct when reviewing an agentic workflow decision?
- Can agents themselves read and rely on tamper-evident process records?
- Can an auditor verify environment state without trusting the executor's self-report?
- How do agent sequences violate system constraints despite individual permissibility?
- What makes a coordination episode revisable under agent intrusion?
- How should task authority constraints apply across multiple coordinated executions?
- What distinguishes sanctioned coordination from intrusion in multi-agent systems?
Related concepts in this collection 7
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can one compromised agent corrupt an entire multi-agent network?
Explores whether a single biased agent can spread behavioral corruption through ordinary messages to downstream agents without any direct adversarial access. Matters because it reveals a previously unknown vulnerability in how multi-agent systems communicate.
both attack MAS without privileged access, but FLOWSTEER acts at planning time while subliminal injection rides ordinary messages at runtime
-
How do adversarial traps target different layers of AI agents?
As AI agents browse the web, attackers can exploit their perception, reasoning, memory, actions, and coordination in distinct ways. Understanding these attack vectors is crucial for building robust agent defenses.
planning-time steering is a systemic trap that the six-category taxonomy frames structurally
-
Can inspecting generated workflows catch planning-time attacks?
Does examining a workflow after it's created catch attacks that corrupt the planning signals upstream? This matters because if contamination enters earlier, downstream inspection might miss malicious intent laundered into legitimate-looking structure.
extends: the defensive corollary — because contamination enters at workflow formation, workflow-inspecting defenses examine an already-compromised artifact
-
How does a signal's position in a workflow change its influence?
Multi-agent systems may amplify or suppress malicious signals based on where they enter the workflow. Understanding position-dependent propagation could reveal which nodes are most critical to defend.
grounds the propagation mechanism: explains why a planning-time bias spreads, since high-influence positions and sycophantic relay amplify the injected signal downstream
-
Do internal agent hops in pipelines need security monitoring?
Multi-agent systems route data between planner, worker, verifier, and synthesizer components. Current defenses only guard user input at the entry point, leaving inter-agent channels unmonitored—but is this a real vulnerability or does downstream safety suffice?
the planner→worker hop this note describes is one of five inter-agent channels in that inventory
-
Which attack and defense numbers came from filtered backends?
Research figures on multi-agent attacks and defenses may have been measured behind provider filters that silently shaped the results. Understanding which numbers had this filter dependency is critical for interpreting their real-world strength.
the 55 percent figure is one the audit would label with its backend and filter status
-
Can semantic labels on requests prevent malicious propagation through agent networks?
SafeFlow explores whether attaching structured intent labels to root requests and propagating them through multi-agent collaboration graphs can block malicious information flow by restoring context that task fragmentation strips away.
a run-time label rides a graph that this attack reshapes at formation; the SafeFlow excerpt does not say how its graph is built or test it against planning-time steering
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- FLOWSTEER: Prompt-Only Workflow Steering Exposes Planning-Time Vulnerabilities in Multi-Agent LLM Systems
- SafeFlow: Semantic Information-Flow Control for Blocking Malicious Propagation in Multi-Agent Systems
- Towards a Science of Scaling Agent Systems
- SoK: When Safe Agents Fail Together: The Security of Multi Agent LLM Systems
- Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection
- Thought Virus: Viral Misalignment via Subliminal Prompting in Multi-Agent Systems
- SchemeArena: Factorized Stress Testing of Scheming in LLM Agents
- A Practical Guide for Designing, Developing, and Deploying Production-Grade Agentic AI Workflows
Original note title
multi-agent planner-executor systems expose a planning-time attack surface where prompts reshape agent organization without touching infrastructure