Securing Agentic AI: From Per-Action Checks to Trajectory Assurance
Autonomous agents are increasingly used to execute consequential tasks in environments governed by operational constraints, organizational policies, regulatory requirements, and technical standards. Their safety is therefore determined not by the correctness of individual actions, but by whether their overall behavior remains consistent with the rules and invariants of the systems in which they operate. As large language model (LLM)-based agents become more autonomous and increasingly delegate tasks across organizational boundaries, securing them evolves from a single challenge into a broad and interconnected landscape spanning the entire agentic stack. At the single-agent level, untrusted inputs through prompts, memory, retrieved knowledge, and tool interfaces create attack surfaces. In multi-agent settings, delegation and communication introduce challenges related to identity, trust, capability control, and decision transparency, while the underlying model routing and execution control plane remains vulnerable to manipulation and to unverified model provenance. Perhaps the most fundamental challenge is behavioral containment: sequences of individually permissible actions may collectively violate system-level constraints and safety invariants.
Introduction. Autonomous agents powered by large language models (LLMs) are increasingly entrusted with critical tasks across enterprise and cyber-physical environments, including healthcare, finance, telecommunications, and critical infrastructure. Unlike traditional AI systems that primarily provide recommendations, these agents plan, reason, invoke tools, interact with external systems, and increasingly collaborate with other agents to accomplish complex objectives. As a result, a single agent may issue thousands of tool calls while a human operator reviews only a handful of decisions. In many of these deployments, agent behavior is constrained by organizational policies, regulatory requirements, or technical standards that specify not only which individual actions are permitted, but also the behavioral envelope within which an entire sequence of actions must remain. Figure 1 illustrates the core components and interaction flows of such an agent. Also agents rarely operate in isolation.
Discussion / Conclusion. Securing autonomous agents is not a collection of isolated challenges but a systems problem spanning the entire agentic stack, from reasoning, memory, and tool use to multi-agent collaboration, model routing, behavioral containment, and ecosystem governance. Across these directions, a common theme emerges: security is fundamentally a property of an agent’s behavior over time and across interacting components, rather than of individual actions in isolation. This calls for a shift from advisory guidance and stateless guardrails to verifiable behavioral invariants and from per-action checks to reasoning about composed, stateful, multi-party behavior. As agents become increasingly autonomous and interconnected, developing such guarantees will be essential for building trustworthy agentic systems and represents a critical research agenda for the security community.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can validator consensus certify semantic correctness beyond agreement? How can infrastructure records verify actual agent behavior? Can single-point security defenses protect multi-agent systems from multi-step attacks? How do coordinated agents balance protocol compliance with reward maximization?- How do compliance concerns drive regulatory scope beyond the stated intent?
- When can the same action count as sanctioned or unsanctioned depending on policy?
- Should agents escalate when facing two equally valid interpretations of a rule?
- How do agent sequences violate system constraints despite individual permissibility?
- Where does an agent's risk come from across its components and sequence?
- What makes violations unavailable rather than merely unchosen in agent architecture?
- How can safety assurance cover whole trajectories at scale?
- What makes uniform bounds the right choice for safety boundaries?
- How do policies distinguish individual action rules from sequence-level constraints?