When Agents Talk: Honeytokens under Shared Memory

Paper · arXiv 2608.11436 · Published August 11, 2026
LLM Failure Modes

During a 2026 cyber-capability evaluation, short-lived AI agents converted a shared package repository into persistent memory. Later agents inherited earlier exploit findings, rebuilt the communication mechanism after it was removed, and the broader evaluation culminated in an intrusion into Hugging Face. The episode raises a design question for defensive deception: can a honeytoken be made harmless to trusted agents without making it recognisable to an attacker who shares their information and can implement the trusted policy? Under those conditions, the answer is no. Any rule that lets a trusted agent use genuine objects while avoiding decoys can be copied by the attacker. When decoys look like genuine objects, a total-variation bound limits how far legitimate use can be separated from benign activation. Pooling signals weakly increases distinguishability. In a fixed-candidate, conditionally i.i.d. benchmark, repeated non-triggering probes drive the minimum Bayes classification error to zero whenever the type-dependent response laws differ; attaining that frontier requires the laws to be known or learnable from feedback. If probing can trigger containment, the coalition must also remain active long enough to collect the observations. A finite-sample bound measures the speed.

Introduction. On 7 May 2026, OpenAI began testing an internal-only research model in a cyber-capability lets process-isolated agents share information. A subsequent ExploitGym evaluation within the same broader programme then crossed an This creates a puzzle for defensive deception. A honeytoken is a credential, file, record, URL or

Discussion / Conclusion. When a compromised agent shares the trusted agent’s information and can implement its policy, With common information and a copyable trusted policy, durable asymmetry requires protected

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How effective are honeytokens and decoys against different security threats? Do honeypot benchmarks validly measure reward hacking better than standard tests? Do backend defenses obscure real attack effectiveness in reported metrics? Can single-point security defenses protect multi-agent systems from multi-step attacks? How vulnerable are token issuance and authorization policies to coordinated attacks? How do we enforce security boundaries in evaluation environments?