SYNTHESIS NOTE
Topics›Agent Harness›this note

How should we measure gains from automatic harness evolution?

Harness evolution itself runs a search loop, so reported improvements might come from more search rather than better design. What's the right way to measure whether the harness itself actually improved?

Synthesis note · 2026-07-17 · sourced from Agent Harness

Automatic harness evolution improves an agent's prompts, tools, memory, and control logic by repeatedly evaluating candidate harnesses against benchmark tasks and revising them. But that revision loop is itself a search procedure — it spends feedback and inference budget hunting for configurations that score well. So when an evolved harness beats the baseline, the gain is confounded: it could come from a genuinely better harness design, or simply from having run more search than the comparison.

The methodological fix is to hold the budget constant. Compare harness evolution against a simple task-level test-time-search baseline given the same feedback signal and the same inference budget. Only the residual — improvement beyond what equal search buys — is attributable to harness design. This is the harness instantiation of a broader eval discipline: because Does a single benchmark score actually predict agent readiness?, a single headline number under uncontrolled budget misleads.

There's a second confound the same paper names: when the evolution search and the final evaluation share one benchmark, reported gains risk overfitting to that task set, so held-out tasks are needed to show the discovered harness generalizes rather than memorizes the test. Together these two controls — matched-budget baselines and held-out evaluation — are what separate "the harness got better" from "we searched harder on the test."

Inquiring lines that read this note 31

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What fundamental constraints limit how effectively agents can improve themselves? How does harness optimization generalize across different model architectures and domains? How can evolutionary algorithms maintain diversity during solution search? How do agent-learned skills transfer and improve across different tasks? How can infrastructure records verify actual agent behavior? What should agent evaluation prioritize to reveal reliable behavior?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 88 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

harness-evolution gains must be measured against test-time search baselines under matched budgets to attribute them to design