SafeFlow: Semantic Information-Flow Control for Blocking Malicious Propagation in Multi-Agent Systems
Multi-agent systems improve capability through task decomposition and role specialization, but these same mechanisms introduce an important safety blind spot: a harmful objective can be fragmented into locally plausible subtasks, allowing malicious intent to evade detection by any single agent. This is a growing social-impact challenge: systems handling sensitive information or consequential tools can turn routine delegation into unauthorized disclosure or unsafe action. We argue that this failure mode is better understood as a semantic information-flow problem than as a single-turn prompt classification task. To address this, we propose SafeFlow, a defense framework for multi-agent systems that formalizes malicious cross-agent propagation as a semantic informationflow problem. SafeFlow attaches structured semantic taints to root requests, propagates them through a dynamic collaboration graph, and performs workflow-level validation to reconstruct the global risk context before irreversible actions are committed.
Introduction. Multi-agent systems increasingly coordinate language models as planners (Yao et al. 2023), tool users (Schick et al. 2023), and role-specialized collaborators (Wu et al. 2023; Hong et al. 2024; Li et al. 2023), with increasingly capable frontier models further accelerating this shift (OpenAI 2023). This creates an insufficiently addressed social safety-andprivacy challenge: planner decisions, inter-agent messages, and tool-side effects jointly determine system behavior. Failures involving sensitive information or consequential tools can affect people and organizations that neither authored the prompt nor observe the resulting workflow. In these systems, harmful behavior often emerges compositionally. A malicious objective can fragment into locally plausible subtasks: one agent retrieves sensitive content, another rewrites it, and a third transmits it, such that no individual step looks overtly malicious, yet the composed workflow realizes exfiltration or policy override.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How does misalignment propagate through agent communication networks?- How do shared state and message propagation transfer failure across agent boundaries?
- What happens when planning signals get contaminated before reaching a downstream agent?
- How does workflow position amplify or suppress malicious signals?
- Can agents rebuild communication channels after removal?
- What interventions prove causation in multi-agent message propagation studies?
- Does anchoring reach communication through unauthorized channels?
- How do unmonitored channels between pipeline agents enable security gaps?
- Can input-boundary defenses guard unmonitored channels between agent hops?
- Why are unmonitored channels between agents a safety risk?
- What makes agent-to-agent messages in multi-agent systems vulnerable to exploitation?
- Can mixed-authorship traces from multi-agent pipelines be monitored reliably?
- What makes unmonitored channels between agents safety-critical?
- How do agent-to-agent messages bypass defenses on downstream principals?
- How do organizations safely retain and control access to committed content?
- What schema do SafeFlow's structured taints use to carry sensitivity information?
- Can SafeFlow distinguish benign uses of sensitive material from actual exfiltration?
- How does workflow-level validation reconstruct risk context from coarse request-level taints?
- How is ground truth defined for labeling harmful outcomes in agent monitoring?
- How does semantic taint survive paraphrase across agent hops?
- How does taint propagation track risk along delegation paths?