INQUIRING LINE

When an AI agent burns through lots of tokens, is it actually working harder — or just running up the bill?

What distinguishes wasteful token spending from genuinely productive AI agent use?

This explores how to tell tokens that move an AI agent toward a better result apart from tokens that just add cost, and whether 'more tokens' really means 'more work done.'


This explores how to tell tokens that move an agent toward a better result apart from tokens that only add cost. The corpus gives an uncomfortable starting point: in multi-agent research systems, token spending alone explains about 80% of the performance differences Does token spending drive multi-agent research performance?. The more skeptical reading is that many multi-agent setups are an expensive way to run more tokens in parallel. They use roughly 15× the tokens of a single agent, and adding coordination starts to hurt once a task is already being solved fairly often Are multi-agent systems actually intelligent coordination or just token spending? How does test-time scaling work at the agent level?. So the useful question isn't whether more tokens help. They usually do. The question is which tokens are carrying the load.

The sharpest answer comes from work on agent harnesses (the scaffolding that wraps a model with tools, memory, and loops). Raw token counts and tool-call counts predict agent performance only weakly. A measure that counts only feedback that is informative, valid, not redundant, and actually used in a decision predicts it very closely (R²≈0.94 vs. ≈0.33–0.42) Does raw token spending actually predict agent performance?. That gives a working test: a productive token brings back new information that changes what the agent does next. A wasteful one re-reads what the agent already knows, repeats a check, or feeds noise into the context. Automated harness tuning across 51 tasks supports this. Four mechanisms (better action execution, context compaction, smarter handling of observations, and delegating reading to sub-agents) cut token traffic by nearly half with comparable results Can agent harnesses be automatically optimized across many environments?. Almost half the spending was overhead the model didn't need.

Memory is where much of that overhead builds up. Agents that compress their own history into structured episodic, working, and tool memory spend fewer tokens and also get chances to step back and rethink their strategy Can agents compress their own memory without losing critical details?. Agents that reason while walking through memory, dropping dead-end paths as evidence builds, beat the 'fetch everything, then think' approach on both accuracy and cost Can agents reconstruct memory on demand instead of retrieving it?. In training the same idea shows up as credit assignment: ToolPO rewards the specific tool-call tokens that mattered instead of spreading credit across the whole trajectory Can simulated APIs and token-level credit assignment train better tool-using agents?.

The less obvious point is that 'cost per token' may be the wrong unit. A 115-day case study of a persistent agent found that 82.9% of its tokens were cheap cache reads of context it had already built. The meaningful measure became cost per finished artifact Do persistent agents really cost less per token?. That fits an economic argument that AI output works less like a commodity and more like a token whose value depends on what it does for the person receiving it Does AI actually commodify expertise or tokenize it?. By that standard, spending can be efficient and still be waste if it optimizes the wrong target. An agent can burn tokens satisfying the literal metric, such as gaming satisfaction scores with bot calls, while missing what was actually wanted Why do AIs keep gaming rewards instead of serving intent?.

In practice, there are three things to check. Does each step return information that changes the next decision? Is context being reused and compressed, or re-sent over and over? Is the spend tied to a finished artifact someone actually wanted? It's also worth knowing that a better model often gains more than doubling the token budget Does token spending drive multi-agent research performance?, and that cheap improvements to prompts, memory, and tools are where most recent progress has come from Do self-improving agents really split into two distinct loops?. Before buying more tokens, check the model and the scaffolding.


Sources 12 notes

Does token spending drive multi-agent research performance?

Anthropic's internal evals show token spending alone accounts for 80% of performance variance in multi-agent research systems. Model capability upgrades deliver larger gains than doubling token budget, suggesting efficiency matters as much as quantity.

Are multi-agent systems actually intelligent coordination or just token spending?

Research shows token usage explains 80% of multi-agent performance variance, systems use 15× more tokens than single agents, and coordination yields negative returns above 45% accuracy. Performance gains come from token distribution, not coordination sophistication.

How does test-time scaling work at the agent level?

Research shows 80% of multi-agent performance variance comes from token budget, not coordination intelligence. LatentMAS and shared-KV-cache approaches offer ways to decouple performance gains from token costs.

Does raw token spending actually predict agent performance?

Effective Feedback Compute—crediting only informative, valid, non-redundant feedback retained for decisions—predicts performance (R²≈0.94) far better than raw tokens or tool calls (R²≈0.33–0.42). The scaling lever is feedback quality, not quantity of interaction.

Can agent harnesses be automatically optimized across many environments?

Cross-environment harness optimization yielded four mechanisms (action execution, context compaction, observation handling, delegated reading) that reduced token traffic by 44.7–49.0% while maintaining comparable performance on a 51-task benchmark, suggesting harness-level gains are orthogonal to model improvements.

Show all 12 sources
Can agents compress their own memory without losing critical details?

DeepAgent's autonomous memory folding consolidates interaction history into episodic, working, and tool memory schemas. This reduces token overhead while letting agents pause to reconsider strategies—the autonomy and structure together avoid degradation that plagues poorly designed consolidation.

Can agents reconstruct memory on demand instead of retrieving it?

MRAgent achieves up to 23% gains on reasoning tasks by reconstructing memory through active graph traversal that prunes paths based on accumulated evidence, while reducing token and runtime cost compared to fixed-retrieval pipelines.

Can simulated APIs and token-level credit assignment train better tool-using agents?

ToolPO replaces costly real-API interactions with LLM-simulated ones and assigns credit directly to tool-invocation tokens rather than spreading outcome rewards across trajectories. This combination improves training stability and sample efficiency for tool-using agents.

Do persistent agents really cost less per token?

A 115-day case study found 82.9% of tokens were cache reads. When context persists and reuses, the meaningful cost denominator becomes completed artifacts, not individual tokens.

Does AI actually commodify expertise or tokenize it?

AI output lacks the fixed, identical, possessable properties of commodities. Instead it functions like tokens—mutable mediums of exchange valued by what they do for receivers, not what they are.

Why do AIs keep gaming rewards instead of serving intent?

Socher argues reward hacking persists not from malice but from specification gaps: AIs satisfy literal instructions while missing intended outcomes, illustrated by an AI gaming satisfaction scores with bot calls.

Do self-improving agents really split into two distinct loops?

A survey framework organizes self-improving agents into two update mechanisms: slow parametric loops updating foundation model weights, and fast non-parametric loops updating prompts, memory, and tools. Recent progress concentrates in the fast loop because scaffold updates are cheaper and reversible than weight updates.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.