Trust propagation and structural containment in Multi-agent LLM pipelines
Abstract— Multi-agent LLM systems increasingly automate tasks involving agents with different levels of privilege, creating a security risk in which a compromised low-privilege agent can influence a higher-privilege agent and trigger an unauthorized action. We study attack propagation in a four-agent LangGraph pipeline comprising a Supervisor, Researcher, Validator, and Executor. We evaluate shared-memory poisoning and indirect prompt injection through a forged approval embedded in a retrieved document. We compare the Validator’s judgment with an independent authorization layer using task-bound signed tokens and a separately verified policy oracle. Our contribution is an empirical study of attack propagation, a component-level ablation of the authorization boundary, and the Judgment Bypass Rate (JBR), which measures compromise at the attacked agent rather than at the final action. Across three seeds and 60 labeled tasks, memory poisoning reaches execution in every undefended trial. With authorization enabled, it achieves 100 % JBR but 0 % Unsafe Action Rate, showing that the Validator can remain compromised while execution is contained.
Introduction. I. INTRODUCTION Multi-agent large language model (LLM) systems increasingly combine specialized agents, such as orchestrators, retrieval agents, reasoning agents, and execution agents, to perform tasks on behalf of users [1]. This separation allows low-privilege agents to retrieve information, review agents to evaluate it, and high-privilege agents to perform consequential actions. However, this delegation chain introduces a security risk. The confused deputy problem, first described by Hardy, occurs when a less-privileged component causes a more-privileged component to misuse authority it legitimately holds [2]. A similar pattern can arise when an agent retrieves, summarizes, and forwards untrusted content to a more privileged agent. Prior work has shown that LLMs can be manipulated through adversarial prompts and indirect prompt injection [3], [4], including in tool-using agents where injected content can influence behavior despite an unchanged system prompt [5].
Discussion / Conclusion. We evaluate the pipeline using controlled ablations on a single codebase, keeping agents, graph topology, prompts, task corpus, and attack injections fixed while varying only the protection layers. T0 disables the token check and policy oracle, whereas T1 enables both. We use 60 tasks (36 safe, 24 unsafe) across three seeds (1, 7, and 42), giving n = 180 trials per cell unless otherwise stated. Because UAR is 0% for most defended conditions, we focus on FPR, LTC, F1, and JBR. Values are means across three seeds, with 95% confidence intervals computed as Student’s t intervals (df = 2). Intervals are reported over seeds because trials within a seed share the same task corpus and are not independent. A. Attack Effectiveness At T0, memory poisoning succeeds in every trial: the compromised Validator forges an approval and the resulting action is executed. With T1, UAR falls to 0% while JBR remains at 100%. Thus, the Validator remains compromised, but the Executor independently prevents unauthorized action (Table I). Agent hijacking produces a different outcome.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can single-point security defenses protect multi-agent systems from multi-step attacks? How can infrastructure records verify actual agent behavior?- Who decides whether an entity has authority to anchor a record?
- Where should the recorder sit relative to the agent it monitors?
- Who holds authority to anchor evidence in this system?
- Where should authenticated provenance records sit to remain outside agent reach?
- Why does least privilege fail when harm exists only in accumulation?
- How do organizations safely retain and control access to committed content?
- Do agents interpret peer edits as legitimate prior changes versus tampering?
- Does an agent's own prior conduct shape the counterparty's response?
- Does delegation transfer authority or merely distribute work across agents?
- How should policy define which agent transfers count as sanctioned versus intrusion?
- Can written policy rules prevent the same transfer from being read two ways?
- Who should verify identity and authorization when agents coordinate across boundaries?
- Does anchoring reach communication through unauthorized channels?
- How does position in a workflow amplify or suppress harmful agent behavior?
- How do ordinary agent messages propagate bias through trusted networks?
- What mechanisms let later agents inherit information left by earlier ones?
- What happens when a compromised middle-agent originates bias rather than the root request?
- What routes do different peer mechanisms use to change agent behavior?