INQUIRING LINE

Tight on tokens? Managing an AI's working memory can cut its token use roughly in half without hurting performance.

How much does context management benefit tasks under tight token budgets?

This explores whether managing what goes into a model's context (compacting it, delegating parts of the work, keeping task state outside the context) still pays off when you can't afford to spend many tokens, and how big that payoff is.


This explores whether managing what goes into a model's context still pays off when tokens are scarce, and how much. The corpus doesn't contain a clean experiment that holds the token budget fixed and then switches context management on and off. What it does have points one way: context management lets you get the same results with roughly half the tokens, and sometimes it gets better results at the same token cost. The clearest number comes from automated tuning of agent harnesses (the code wrapped around a model) across many environments. That process found four mechanisms: executing actions directly, compacting context, handling tool observations more economically, and handing reading tasks off to helpers. Together they cut token traffic by 44.7–49.0% while keeping performance about the same on a 51-task benchmark Can agent harnesses be automatically optimized across many environments?. Those gains came from the harness, not from a better model, so they stack on top of model upgrades.

This matters because spending more tokens is usually the way systems get better. Anthropic's internal evaluations found that token spending alone explains about 80% of the performance differences between multi-agent research systems Does token spending drive multi-agent research performance?. When the budget is capped, that lever is gone, and making each token count is the main thing left to improve. Several notes show this efficiency turning into real gains. Training a model to hand subtasks to subagents and take back only their summaries beat passive compression, and it let a 30B model match much larger ones Can delegation teach models to manage context more actively?. Explicit algorithms that show each LLM call only the context its current step needs get around context-window limits by design Can algorithms control LLM reasoning better than LLMs alone?. Memory that is rebuilt by walking a graph and pruning dead ends as evidence comes in improved reasoning by up to 23% and also cost fewer tokens than a fixed retrieve-then-reason pipeline Can agents reconstruct memory on demand instead of retrieving it?.

The biggest single jump in the corpus treats the problem as one of tracking task state rather than trimming tokens. Keeping a checked record of task progress outside the model's context, and confirming progress against the environment instead of trusting the model's own claims, raised one model from 51.8% to 80.7% on a long-horizon benchmark Can task state management alone improve long-horizon agent performance?. A related note argues that agents in long workflows fail because of weak control over what gets written to memory, not because they lack knowledge. Replaying the whole transcript lets errors and drifting constraints pile up Can agents fail from weak memory control rather than missing knowledge?. The takeaway is that a tight budget forces you to decide what the model keeps, and making that decision well can improve accuracy, not just cut cost.

Two framings may change how you think about 'budget' itself. First, the scarce resource may be compute rather than space. One line of work finds the long-context bottleneck is the compute needed to absorb context that has been pushed out of the window into the model's internal state, and performance rises with more of these consolidation passes Is long-context bottleneck really about memory or compute?. A related approach structures reasoning as nested subtask trees that prune their own memory cache (the KV cache) as they go, so the model can keep reasoning past its window limit Can recursive subtask trees overcome context window limits?. Second, the token may be the wrong unit of cost. In a 115-day case study, 82.9% of tokens were cache reads, which pushes the meaningful cost measure toward finished work products rather than raw tokens Do persistent agents really cost less per token?. Under that view, good context management also means arranging context so it can be reused cheaply.

There is one limit. Context management makes better use of a budget, but it can't fix everything. Some tasks need different kinds of expertise, parallel work and independent checking, which no single agent loop can organize however well it handles its context Do single agents always hit organizational limits?. A practical partner technique is to route the routine subtasks to small models at 10–30× lower cost, which saves budget in a different way Can small language models handle most agent tasks?.


Sources 12 notes

Can agent harnesses be automatically optimized across many environments?

Cross-environment harness optimization yielded four mechanisms (action execution, context compaction, observation handling, delegated reading) that reduced token traffic by 44.7–49.0% while maintaining comparable performance on a 51-task benchmark, suggesting harness-level gains are orthogonal to model improvements.

Does token spending drive multi-agent research performance?

Anthropic's internal evals show token spending alone accounts for 80% of performance variance in multi-agent research systems. Model capability upgrades deliver larger gains than doubling token budget, suggesting efficiency matters as much as quantity.

Can delegation teach models to manage context more actively?

SearchSwarm shows that training models to delegate subtasks and integrate summarized results beats passive compression, with a 30B model matching much larger ones. Critically, the delegation skill transfers to single-agent tasks, suggesting it teaches disciplined decomposition and evidence grounding, not just orchestration.

Can algorithms control LLM reasoning better than LLMs alone?

LLM Programs embed LLMs within explicit algorithms that manage control flow and state, presenting only step-specific context to each LLM call. This information hiding addresses capability and context window limits while treating complex reasoning as modular, debuggable sub-tasks.

Can agents reconstruct memory on demand instead of retrieving it?

MRAgent achieves up to 23% gains on reasoning tasks by reconstructing memory through active graph traversal that prunes paths based on accumulated evidence, while reducing token and runtime cost compared to fixed-retrieval pipelines.

Show all 12 sources
Can task state management alone improve long-horizon agent performance?

Separating task state management from execution, using independent environment audits instead of trusting executor claims, improved Qwen 3.7-Plus from 51.8% to 80.7% on WeaveBench. The same model-harness pair showed consistent gains across multiple benchmarks and task types.

Can agents fail from weak memory control rather than missing knowledge?

Agent performance degrades in long workflows because transcript replay and retrieval-based memory lack gating mechanisms. A bounded, schema-governed committed state that separates artifact recall from permanent memory write prevents error accumulation and constraint drift.

Is long-context bottleneck really about memory or compute?

Research shows the bottleneck is not memory capacity but the compute required to consolidate evicted context into fast weights during offline sleep phases. Performance improves with more consolidation passes, following a test-time scaling pattern on harder reasoning tasks.

Can recursive subtask trees overcome context window limits?

The Thread Inference Model demonstrates that reasoning structured as recursive subtask trees with rule-based KV cache pruning sustains accurate reasoning beyond context limits, even when manipulating 90% of the cache. This enables single models to replace multi-agent systems by handling full recursive reasoning internally.

Do persistent agents really cost less per token?

A 115-day case study found 82.9% of tokens were cache reads. When context persists and reuses, the meaningful cost denominator becomes completed artifacts, not individual tokens.

Do single agents always hit organizational limits?

Research shows that real-world tasks requiring heterogeneous expertise, parallel execution, and independent verification exceed what any single agent loop can organize. Graph-based system abstractions are needed to distribute intelligence across specialized agents.

Can small language models handle most agent tasks?

SLMs handle the repetitive, well-defined language tasks that constitute most agent work at 10–30× lower cost than LLMs, making heterogeneous architectures (SLMs by default, LLMs selective) the economically rational design pattern.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.