When an AI tunes its own setup, the feedback signal that feels most reliable — task success — turns out to be the most misleading.
What feedback signals matter most during harness evolution search?
This explores which kinds of feedback — task success, held-out evaluation, persistence in the loop, external anchors — actually drive improvement when an agent iteratively rewrites its own scaffolding (prompts, tools, objectives) through a search process.
This explores what actually moves the needle when an agent evolves its own harness — the prompts, tools, and objectives that scaffold its behavior — rather than assuming any feedback loop that produces changes is producing improvement. The corpus's sharpest warning is that the most seductive signal, raw task success, is also the most misleading: harness evolution is itself a search loop, so gains reported against it are confounded with sheer search effort How should we measure gains from automatic harness evolution?. The feedback that matters is the delta *beyond* a matched search budget, measured on held-out tasks — otherwise you're rewarding memorization, not design. This is borne out by inspection: most evolved edits just cache task-specific fixes the agent could have rediscovered in a single rollout, so the signal to distrust is one that improves already-solvable tasks rather than converting failures into successes Do harness edits learn reusable strategies or memorize task fixes?.
If outcome scores are noisy, the corpus points to a more reliable signal hiding underneath them: persistence through the loop. Across 17 frontier models on long-horizon optimization tasks, the dominant predictor of success wasn't the quality of any single edit but whether the agent kept running the benchmark-edit-incorporate cycle instead of terminating early or burning budget unproductively What predicts success in ultra-long-horizon agent tasks?. So during search, a signal worth tracking is whether feedback is being *incorporated* iteration over iteration — not just whether any one iteration scored well. Feeding that loop requires dense, per-step signals rather than a lone scalar, which is why decoupling evaluation into separate benchmark, harness, and environment components matters: it surfaces reward-hacking and failure modes in the trajectory that a single success number hides entirely How can we make reward-hacking visible in agent evaluation?.
The deepest constraint is that self-generated feedback alone runs out. Pure self-improvement stalls on the generation-verification gap and reward hacking, and the methods that actually work smuggle in external anchors — past model versions, third-party judges, user corrections, tool feedback Can models reliably improve themselves without external feedback?. Harness evolution is a self-improvement loop, so the same rule applies: the feedback that keeps it honest comes from outside the agent's own judgment. Tree search offers one way to manufacture such a signal without human labels — MCTS outcomes plus critic models yield dense process-level rewards that rank solution paths by whether they actually succeed Can tree search replace human feedback in LLM training?.
There's also a wrinkle in *who* the feedback is for. The capacity to produce useful harness edits is flat across model tiers, but the capacity to benefit from them follows an inverted U — weak models never invoke the harness, strong ones struggle to follow their own instructions faithfully Do stronger models always evolve harnesses better?. So a signal that works for a mid-tier planner may be inert for others, and how the harness is *represented* changes what feedback is even actionable: reorganizing it around runtime behavior lets a weaker planner localize as well as a stronger model Can explicit behavior maps help weaker planners compete with stronger models?.
Finally, the frontier framing turns feedback from an input into something the search itself produces. Evolutionary inference-time search beats Best-of-N and sequential revision partly because an island model sustains population diversity — the signal it needs is not just 'is this good' but 'is this *different*,' since entropy-collapsed candidates can't be recombined into anything new Can evolutionary search beat sampling and revision at inference time?. That's why training upstream for diversity beats scalar optimization when a model feeds into search Should training maximize diversity when models feed into search?. And a bi-level agent can close the loop entirely by evolving its own objective functions — compiling natural-language goals into executable scoring code — so the feedback signal becomes a first-class thing the search designs rather than inherits Can agents evolve their own objectives during search?. The through-line: the feedback that matters most is external, budget-matched, diversity-aware, and measured by conversion of failures — not the raw task score the loop is happiest to optimize.
Sources 11 notes
Automatic harness evolution must be compared against task-level test-time search under equal feedback and inference budgets. Only gains beyond what matched search achieves are attributable to the harness design itself, not just more computation.
Analysis of evolved harness trajectories shows rational, well-motivated edits across prompt and tool layers, but most persist fixes an agent could rediscover in a single rollout. Gains remain limited because memorized shortcuts cache what's already within reach rather than converting hard failures into successes.
Across 17 frontier models on 36 expert-curated optimization tasks, repeated benchmark-edit-incorporate cycles within a wall-clock budget proved the dominant success predictor. Most models terminated early or burned budget unproductively; Claude Opus 4.6 stood out as persistent.
AgentCompass separates benchmark, harness, and environment into independent components, enabling trajectory analysis that surfaces reward-hacking and other failure modes scalar scores conceal. This architectural change shifts evaluation from opaque final scores to inspectable agent behavior.
Pure self-improvement stalls due to the generation-verification gap, diversity collapse, and reward hacking. Reliable improvement methods succeed by smuggling in external anchors: past model versions, third-party judges, user corrections, or tool feedback.
Show all 11 sources
AlphaLLM uses tree search outcomes and three critic models to derive dense reward signals equivalent to human-labeled feedback. Tree structure naturally ranks solution paths by success, replacing the annotation oracle that standard RLHF requires.
Model capability to produce useful harness edits stays constant across tiers, but capacity to actually benefit from those edits follows an inverted U-shape, peaking in mid-tier models. Weak models fail to invoke harnesses; strong models struggle with faithful instruction-following.
A behavior-to-code mapping representation improved win rates by 10–19 points while reducing planner tokens by 8–13%. Weaker planners using this mapping matched stronger models' code localization across all precision and recall metrics.
Mind Evolution, an evolutionary search strategy using LLM-generated crossover and mutation with island model diversity, solves 98%+ of planning tasks and significantly outperforms best-of-N and sequential revision strategies while working directly in natural language without task formalization.
Vector Policy Optimization trains models to emit varied competent solutions rather than converging to one answer. This unlocks search procedures like evolutionary algorithms to explore and combine modes, solving problems that entropy-collapsed policies cannot reach at all.
SAGA's bi-level architecture closes a feedback loop from optimization results back to goal design by having an outer LLM loop propose new objectives and compile them into code the inner loop can immediately use, enabling objective formulation as part of discovery rather than a fixed input.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
- Rethinking the Evaluation of Harness Evolution for Agents
- Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
- HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
- Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution
- DarwinX: Evolving Agent Harnesses Through Natural Selection
- StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
- RRSI: Regularized Recursive Self-Improvement of Agent Harnesses