Can an external manager handle context for frozen agents?
Exploring whether a separate trained system can effectively manage a frozen agent's context window. This matters because many deployed agents are closed-source and can't be retrained, yet they suffer from context degradation.
Long-horizon agents accumulate context — tool results, intermediate reasoning — until stale content obscures salient evidence, amplifies positional bias, and degrades decisions. Prior fixes put the burden of managing context on the agent itself (agent-side control, or fixed summarization), which requires training the agent and is impractical for closed-source agents, and ignores that different agents need different strategies.
AdaCoM separates the concern entirely: train an external LLM to manage the context of a frozen agent through flexible modification actions and end-to-end RL. The manager prunes stale content while preserving task constraints and progress, improving diverse agents on web-search and deep-research benchmarks and transferring to unseen agents of similar capability.
The most useful finding is a fidelity–reliability trade-off. Agents with higher vanilla ReAct performance benefit from higher-fidelity context preservation — they can use more detail well. Lower-performing agents require more aggressive compression to stay within a reliable reasoning regime. The right amount of context is not a property of the task alone; it is indexed to the agent's own competence. This means context management is not one universal policy but a per-agent calibration — consistent with Does fixed sparsity work for all sequence lengths?, where the optimal budget is also conditional rather than fixed.
Inquiring lines that read this note 44
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How should designers communicate what AI systems truly are and can do?- Why does continuous agent inference differ from human user inference?
- Why is digital context more volatile than conventional software context?
- How do perception and execution gaps limit current AI agent performance?
- How does context engineering bridge human intent and machine understanding?
- Can the same compress-then-act pattern work for agent state memory?
- How does external context control compare to agents managing their own state internally?
- How should agents compress episodic interactions into working memory without accumulation?
- Can externalizing bookkeeping to a stateful harness replace internalized memory control?
- How do memory hygiene and context efficiency trade off in deployed agents?
- Why do agents ignore condensed experience in favor of raw data?
- How should embedding model speed constrain agent memory system design?
- What shapes of memory help frozen agents improve without retraining?
- Can task-agnostic compression of documents remain broadly useful for later queries?
- How do memory hierarchies and compression reduce context management demands?
- Why do weaker agents need more aggressive context compression than stronger ones?
- How does reducing activation precision further extend context length?
- Why does keeping full key-value blocks matter more than compressing them?
- What is the connection between model compression and data compression?
- Does recurrent memory or gist compression work better for ultra-long context?
- Can external managers optimize context better than the model itself?
- Can fixed-size latent states losslessly store arbitrary input context?
- Why is long-context compute spent transforming context into internal state rather than storing it?
- What components of agent scaffolding most impact domain-specific output quality?
- Can harness updates benefit agents equally across all model sizes?
- How should harness scaffolding be treated as a first-class object?
- How can harnesses externalize bookkeeping so models focus on semantic judgment?
- Should optimal context budgets scale with agent competence or task complexity?
- Can context management policies transfer across agents of similar capability levels?
- How much does external context management transfer across similar capability agents?
- Can context management be optimized for an agent without retraining or changing the model?
- Why should consolidation be scheduled offline rather than during forward passes?
- Why does consolidating more state sometimes hurt performance below the no-memory baseline?
- Why does externalized state beat parameter scaling for agent reliability?
- How does externalizing reasoning into harness artifacts improve agent reliability?
- How does structured environment-side state reduce multi-turn agent failure better than transcript replay?
- What mechanism explains why context management prevents overflow failures most?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can agents fail from weak memory control rather than missing knowledge?
As multi-turn agent workflows grow longer, performance degrades—but is this due to insufficient context or poor memory management? This explores whether memory *control* is the real bottleneck.
adjacent solution to the same accumulation problem; ACC commits state internally, AdaCoM manages it externally
-
Can context playbooks prevent knowledge loss during iteration?
When AI systems iteratively refine their instructions and memories, do structured incremental updates better preserve domain knowledge than traditional rewriting? This matters because context degradation undermines long-term agent performance.
both treat context as actively managed rather than passively appended
-
Where does agent reliability actually come from?
Exploring whether LLM agent performance depends on larger models or on thoughtful system design choices like memory, skills, and protocols that shift cognitive work outside the model.
the external manager is harness infrastructure for a frozen model
-
What problems did AIDE2's rewrites actually solve?
AIDE2 autonomously improved its own code over eight days. Did the seven accepted changes target real practitioner challenges in building agentic systems, or did they reflect artifacts of the system's own optimization process?
the same problem class reached from inside the agent: a self-rewriting loop kept "memory mechanisms that compress and manage the agent's growing context" among its seven rewrites; the excerpt does not say what they do or whether they adapt to the agent's reliability
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Learning Agent-Compatible Context Management for Long-Horizon Tasks
- SearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep Research
- RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
- LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents
- From Model Scaling to System Scaling: Scaling the Harness in Agentic AI
- Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems
- Adaptation of Agentic AI
- Are We Ready For An Agent-Native Memory System?
Original note title
context management can be offloaded to a trained external manager for a frozen agent and optimal compression depends on the agent's own reliability