SYNTHESIS NOTE
Topics›Agent Harness›this note

How should we measure agent system performance beyond task success?

Current evaluation metrics collapse agent behavior into a single success score, hiding critical information about how agents operate. What dimensions—trajectory quality, memory use, context efficiency, verification cost—should benchmarks actually measure?

Synthesis note · 2026-05-28 · sourced from Agent Harness

Agent evaluation has inherited the model-centric habit of reducing performance to a single number: final-task success or benchmark accuracy. The "system scaling" framing argues this framing is increasingly inadequate, because agent behavior emerges from the interaction of the foundation model with a memory substrate, a context constructor, a skill-routing layer, an orchestration loop, and a verification-and-governance layer. A one-shot success score collapses all of this into a binary that hides how the agent got there. Two agents with identical task-success rates can differ enormously in how much they spent, how much context they wasted, how clean their memory stayed, and how reliably they verified their own actions.

The proposed alternative is a research agenda for harness-level benchmarks that measure trajectory quality, memory hygiene, context efficiency, communication fidelity, verification cost, and safe evolution over time. The point is that the same model "projected onto different harnesses produce qualitatively different agents" — so evaluation must measure the system, not just the model. The counterpoint is that multi-dimensional metrics are harder to optimize and compare, and task success remains the outcome users ultimately care about. But success-only scores create false confidence in deployment readiness. This matters because it tells builders what to instrument: the process, not only the outcome.

Inquiring lines that read this note 130

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do capability benchmark scores systematically misrepresent true model abilities? Should agents decouple planning from perception grounding for better performance? Why do agents falsely report success on failed tasks? What execution architectures enable agents to most effectively use tools? How does the generation-verification gap limit what we can measure about AI reasoning? How do agent-learned skills transfer and improve across different tasks? When should work require human-AI partnership versus full automation? How do standardized protocols improve multi-agent coordination and reliability? What should agent evaluation prioritize to reveal reliable behavior? Do multi-agent systems introduce security vulnerabilities that single-agent architectures avoid? Can harness architecture and protocols provide agent reliability without model scaling? Do reasoning benchmarks predict model performance in long-horizon workflows? When do multi-agent systems outperform single frontier models? Why do standard benchmarks fail to predict agent deployment success? When do multi-agent systems provide sufficient quality returns on token investment? How does evaluation scope and dimensionality affect what we measure? How should test-time compute scaling work in agentic systems? How does harness optimization generalize across different model architectures and domains? What trajectory-level metrics beyond task success best evaluate agent performance? How should agent systems validate and persist generated code artifacts? How can infrastructure records verify actual agent behavior? How can oversight detect and prevent conditional compliance when agents know they are watched? What fundamental constraints limit how effectively agents can improve themselves? Can multi-agent systems avoid converging on false agreement without deliberation? How do we enforce security boundaries in evaluation environments? Can causal models help detect and locate hidden sandbagging in AI? Do backend defenses obscure real attack effectiveness in reported metrics? What determines whether deployed AI systems can actually be stopped in practice? How should agents manage memory granularity to improve long-term performance?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
24 direct connections · 195 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

agent evaluation must move beyond one-shot task success to trajectory quality memory hygiene context efficiency and verification cost