EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents

Paper · arXiv 2609.05903 · Published September 5, 2026
Agent Harness

Large Language Model (LLM) Agents are turning language into real-world effects. They should remain safe against both indirect prompt injections and direct harmful requests. System-level safety harnesses provide an additional enforcement layer in addition to model-level solutions, but existing harness designs are typically built once by experts and applied across heterogeneous models and domains. The effective defense is inherently deployment-dependent: models differ in how much external enforcement they need before utility starts to drop, while domains differ in which effects, state, and action sequences must be governed. A harness strict enough for one model over-blocks another, and a policy general enough to transfer across domains can miss the safety relations of the application. We present EvoSafeHarness, a safety-specific harness optimization framework that automatically synthesizes a deployable harness for a frozen model in a target domain.

Introduction. Language-model agents are moving from demonstration to deployment. As they gain access to sensitive data, financial accounts, production systems, and external services, safety becomes an operational requirement. A chatbot failure may end in an undesirable response; an agent failure can result in a transferred payment, a leaked credential, deleted production data, or a persistent shell process. The unit of safety has expanded from a single utterance to an entire action trajectory, and the consequences of failure have expanded with it. What makes agent safety qualitatively harder is that harmful instructions can enter through two channels that cross different security boundaries. In an indirect prompt injection attack [Greshake et al., 2023, Liu et al., 2024, Perez and Ribeiro, 2022], an adversary embeds instructions in external content—such as an email, a web page, or a document—that the agent must consume as data. If the agent treats this content as authoritative, it may execute actions that the user never requested.

Discussion / Conclusion. A harness should not be copied unchanged across deployments. The domain tells us what can go wrong and what state must be protected; the victim model and runtime tell us which checks are needed, where they must run, and what history they must remember. The domain supplies the tool semantics, trusted boundaries, and threat contract that determine the relation to protect and the state it requires. The victim and runtime determine its realization: semantic or deterministic, cached or recomputed, and enforced before or after execution. This is why EvoSafeHarness optimizes a deployment instead of selecting one portable guardrail. 8 Conclusion

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Can single-point security defenses protect multi-agent systems from multi-step attacks? Do multi-agent systems introduce security vulnerabilities that single-agent architectures avoid? How should agent systems validate and persist generated code artifacts? Why do locally safe actions create system-level safety gaps? How does misalignment propagate through agent communication networks? How can infrastructure records verify actual agent behavior? How do coordinated agents balance protocol compliance with reward maximization? How can we detect and prevent harm propagation through multi-agent delegation workflows? How vulnerable are token issuance and authorization policies to coordinated attacks? How do false presuppositions and sycophancy drive persistent false beliefs in models? Can local safety checks guarantee system-level behavioral safety?