SYNTHESIS NOTE
Topics›Evaluations›this note

How should we evaluate agent behavior beyond final answers?

As AI systems move from single-response tasks to multi-step interactions, what evidence should evaluation focus on? This explores whether scoring interaction trajectories alongside process quality, recovery, and coordination reveals system capabilities that final-answer metrics miss.

Synthesis note · 2026-05-28 · sourced from Evaluations

If evaluation is the map E: X → Y from admissible evidence to judgments, then the shift to agentic systems changes both terms in a parallel, recurring way. On the evidence side (X), the unit expands from a single final response to a full interaction-generated trajectory — the sequence of states, actions, tool calls, and environment responses produced as the system acts in closed loop. On the procedure side (E), final correctness is no longer sufficient; the evaluator must additionally score process quality, recoverability (can the agent get back on track after an error?), coordination (across tools, environments, other agents), robustness, efficiency, and system-level performance.

This is a pattern, not a single metric, because the same expansion recurs across otherwise unrelated agent benchmarks. T-Eval scores whether each predicted tool call matches the expected one; AgentBoard's Progress Rate compares the actual trajectory against the expected trajectory; multi-agent frameworks score collaborative efficiency and how well agents distribute tasks dynamically. Each is an instance of "stop scoring the endpoint, start scoring the path." The trajectory becomes the evidence, and the qualities that only exist over time — recovery, coordination, partial progress — become the things judged.

Why it matters: this reframes a scattered set of agent metrics as a coherent move. Once you see process-recoverability-coordination scoring as the trajectory-level analogue of final-answer scoring, you can ask the design-science questions — which artifacts to admit, how to map them to judgments — systematically rather than benchmark by benchmark. The counterpoint: richer evidence is also noisier and harder to standardize, which is precisely why the expansion creates new evaluation challenges rather than dissolving the old ones.

Inquiring lines that read this note 62

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Why do agents falsely report success on failed tasks? How do evaluation practices shape which failures stay visible? What should agent evaluation prioritize to reveal reliable behavior? How does the generation-verification gap limit what we can measure about AI reasoning? How do standardized protocols improve multi-agent coordination and reliability? Can harness architecture and protocols provide agent reliability without model scaling? When do multi-agent systems outperform single frontier models? What trajectory-level metrics beyond task success best evaluate agent performance? How can infrastructure records verify actual agent behavior? What fundamental constraints limit how effectively agents can improve themselves? Why do standard benchmarks fail to predict agent deployment success? Do honeypot benchmarks validly measure reward hacking better than standard tests? How should agents manage memory granularity to improve long-term performance? What determines whether deployed AI systems can actually be stopped in practice? How do coordinated agents balance protocol compliance with reward maximization? Can single-point security defenses protect multi-agent systems from multi-step attacks? How does evaluation scope and dimensionality affect what we measure? What prevents conversational agents from taking initiative in dialogue? How can humans maintain meaningful oversight as AI systems become increasingly autonomous and complex? Does AI assistance promote real skill development or substitute for independent learning? When should work require human-AI partnership versus full automation? Should agents decouple planning from perception grounding for better performance? Do reasoning benchmarks predict model performance in long-horizon workflows? What safeguards enable trustworthy AI-assisted scientific peer review at scale? How do agent-learned skills transfer and improve across different tasks? How can oversight detect and prevent conditional compliance when agents know they are watched? Why do locally safe actions create system-level safety gaps?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
18 direct connections · 186 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

agent evaluation expands evidence from final responses to interaction trajectories scoring process recoverability and coordination