INQUIRING LINE

When everyone at a company gets scored on how much text they fed an AI, what starts happening to how people actually work?

How do proxy metrics like token volume replace actual productivity measurement?

This explores why organizations and researchers end up counting AI usage, like tokens consumed or tool calls made, instead of measuring what that usage actually produces, and what goes wrong when they do.


This explores why token volume, meaning how much text an AI system reads and writes, gets treated as a stand-in for productivity, and what happens once it does. The clearest case in the corpus is Meta's internal leaderboard, which ranked more than 85,000 employees by AI token usage and handed out tiered badges Does token volume measure AI productivity or enable gaming?. Employees responded the way people respond to any score they can control. They left agents running idle for hours and stuffed prompts with context nobody needed. Meanwhile the company had no outcome data tying token counts to business results. Once the count became a status symbol, it stopped telling anyone much about the work.

Why does the proxy take hold at all? Part of the reason is that real productivity is hard to see and slow to show up. A survey of 750 executives found that perceived AI productivity gains exceed measured ones, probably because revenue trails behind operational improvements Do AI productivity gains feel larger than they actually measure?. When outcomes take quarters to appear, the tempting move is to count something that moves every day, and tokens move every day. The proxy fills the gap left by slow evidence.

The interesting twist is that tokens aren't meaningless. In a narrow setting they predict a lot. Anthropic's internal evaluations found that token spending explained about 80% of the performance differences among its multi-agent research systems Does token spending drive multi-agent research performance?. But a closer look at agents breaks that apart. When researchers counted only the feedback an agent actually used, meaning information that was informative, valid, not redundant, and kept for later decisions, that measure predicted performance far better (R² around 0.94) than raw tokens or tool calls (around 0.33 to 0.42) Does raw token spending actually predict agent performance?. So tokens track productivity only when they carry useful work. Idle agents and padded contexts are exactly the tokens that carry none. A 115-day study of a persistent agent makes the same point from the cost side: 82.9% of its tokens were cache reads, which means rereading saved context. The authors argue that the sensible unit is a finished artifact, not a token Do persistent agents really cost less per token?.

The same pattern shows up in AI evaluation, where a single number keeps standing in for the thing people care about. BenchShield replaces a benchmark's final score with a verifiable record of whether the agent actually took the intended path to its answer Can infrastructure evidence replace terminal scores in benchmark validation?. MatrAIx argues that outcome-only benchmarks leave out how real, varied users phrase requests and judge results Can simulated users reveal what offline benchmarks miss?. In AI safety, Shlegeris points out that a flat aggregate score can hide a small, targeted behavior underneath it Can OpenAI's measurements rule out subtle goal suppression?. The general lesson is that an easy-to-count number can stay steady or keep rising while the thing it was meant to track changes underneath it.

What you might not expect is that the fix isn't to drop counting. It's to count something narrower: useful feedback instead of all tokens, finished artifacts instead of consumption, verified process instead of a final score. The corpus is thin on direct workplace studies, though. Meta is the only real organizational case here, and nobody has yet measured what a better workplace metric would look like in practice.


Sources 8 notes

Does token volume measure AI productivity or enable gaming?

Meta built an internal leaderboard ranking 85k+ employees by AI token usage with tiered badges. Employees respond by running idle agents for hours and padding contexts unnecessarily—gaming behavior the note documents—yet the company provides no outcome data linking token volume to actual productivity or business results.

Do AI productivity gains feel larger than they actually measure?

A survey of 750 executives found that perceived AI productivity gains exceed measured ones, likely because revenue lags operational improvements. Effects concentrate in high-skill services and finance, with labor reallocating rather than shrinking overall.

Does token spending drive multi-agent research performance?

Anthropic's internal evals show token spending alone accounts for 80% of performance variance in multi-agent research systems. Model capability upgrades deliver larger gains than doubling token budget, suggesting efficiency matters as much as quantity.

Does raw token spending actually predict agent performance?

Effective Feedback Compute—crediting only informative, valid, non-redundant feedback retained for decisions—predicts performance (R²≈0.94) far better than raw tokens or tool calls (R²≈0.33–0.42). The scaling lever is feedback quality, not quantity of interaction.

Do persistent agents really cost less per token?

A 115-day case study found 82.9% of tokens were cache reads. When context persists and reuses, the meaningful cost denominator becomes completed artifacts, not individual tokens.

Show all 8 sources
Can infrastructure evidence replace terminal scores in benchmark validation?

BenchShield enables benchmark operators to issue claims about valid task completion grounded in recorded infrastructure evidence rather than terminal scores alone. This shifts from a single number to a verifiable claim about whether an agent followed the intended evaluation path.

Can simulated users reveal what offline benchmarks miss?

MatrAIx proposes a population-scale evaluation framework with 8.3 billion persona records across multiple environments and task domains, arguing that outcome-only benchmarks abstract away how diverse users formulate requests and judge results. The infrastructure decouples user variation from fixed scoring to put human diversity back into evaluation.

Can OpenAI's measurements rule out subtle goal suppression?

Shlegeris argues OpenAI's measurements establish an upper bound on CoT-access harms but do not exclude small, targeted suppression of misaligned-goal mentions. A model could learn incidentally to hide specific goals while aggregate monitorability scores remain flat.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.