How should we measure gains from automatic harness evolution?
Harness evolution itself runs a search loop, so reported improvements might come from more search rather than better design. What's the right way to measure whether the harness itself actually improved?
Automatic harness evolution improves an agent's prompts, tools, memory, and control logic by repeatedly evaluating candidate harnesses against benchmark tasks and revising them. But that revision loop is itself a search procedure — it spends feedback and inference budget hunting for configurations that score well. So when an evolved harness beats the baseline, the gain is confounded: it could come from a genuinely better harness design, or simply from having run more search than the comparison.
The methodological fix is to hold the budget constant. Compare harness evolution against a simple task-level test-time-search baseline given the same feedback signal and the same inference budget. Only the residual — improvement beyond what equal search buys — is attributable to harness design. This is the harness instantiation of a broader eval discipline: because Does a single benchmark score actually predict agent readiness?, a single headline number under uncontrolled budget misleads.
There's a second confound the same paper names: when the evolution search and the final evaluation share one benchmark, reported gains risk overfitting to that task set, so held-out tasks are needed to show the discovered harness generalizes rather than memorizes the test. Together these two controls — matched-budget baselines and held-out evaluation — are what separate "the harness got better" from "we searched harder on the test."
Inquiring lines that read this note 31
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What fundamental constraints limit how effectively agents can improve themselves? How does harness optimization generalize across different model architectures and domains?- Can harness updates benefit agents equally across all model sizes?
- How should harness scaffolding be treated as a first-class object?
- What makes harnesses more tangled than other types of agent code?
- Why does the harness layer accumulate distributed behaviors over time?
- Why do mid-tier models benefit more from memorized harness shortcuts?
- Can harness evolution be redirected toward distilling transferable procedures instead?
- What persistent failures remain unsolved despite harness evolution efforts?
- Why do evolved harness edits mostly memorize rather than generalize?
- What feedback signals matter most during harness evolution search?
- What makes behavior localization the bottleneck in agent harness evolution?
- What makes a harness a first-class object rather than invisible scaffolding?
- Why do mid-tier models benefit most from memorized harness fixes?
- How much of harness-evolution gain comes from matched test-time search budgets?
- Why do evolved harnesses often fail to generalize beyond their training tasks?
- Why does harness benefit capacity peak at mid-tier models, not frontier scale?
- How does editing the harness layer differ from updating model weights?
- Which foundation model tiers most benefit from harness updates?
- How do evolved harness edits generalize across different benchmark domains?
- What makes a harness low-friction for model strategy?
- Which domains see models exceed human harness design quality?
- Why do useful harness updates often disappear during model evolution?
- How much does executor choice change a harness's actual performance?
- Can mid-tier models benefit more from harness improvements than frontier models?
- How much does harness design contribute to reported model capability scores?
- How much of DarwinX's gain comes from maintaining an archive versus single-lineage search?
- Can harness evolution gains be distinguished from test-time search improvements on matched budgets?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do harness edits learn reusable strategies or memorize task fixes?
When meta-agents evolve harnesses iteratively, do the persisted edits encode transferable procedures that solve new problems, or do they mostly cache shortcuts for already-solvable tasks? This matters because it determines whether harness evolution genuinely expands capability.
the mechanism behind why matched-budget gains stay small
-
Does a single benchmark score actually predict agent readiness?
Single-axis benchmarks rank models by one capability—like task success—but ignore privacy, duration, operating mode, and ecosystem fit. Can one number really capture what matters for deployment?
same eval discipline applied to agent benchmarks generally
-
Can distance alone rank which substrates resist reward hacking?
Does the amount a method changes a model reliably predict how exposed it is to evaluator errors? The paper tests whether a single distance metric can universally order vulnerability across weights, selection, and prompts.
lists optimization budget among the factors that shape exposure to an evaluator's mistakes; matching budgets holds that one factor fixed and leaves the location of a scoring defect free, which that note's constructed illustration shows can change which behavior each method favors
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Rethinking the Evaluation of Harness Evolution for Agents
- HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
- Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
- DarwinX: Evolving Agent Harnesses Through Natural Selection
- ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement
- Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution
- Scaling Laws for Agent Harnesses via Effective Feedback Compute
- RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
Original note title
harness-evolution gains must be measured against test-time search baselines under matched budgets to attribute them to design