Self-improving AI tooling mostly memorizes fixes for tasks it could already do — the genuinely hard failures go unaddressed.
What persistent failures remain unsolved despite harness evolution efforts?
This explores what problems stubbornly persist even after teams try to make AI agent 'harnesses' — the prompts, tools, and scaffolding wrapped around a model — evolve and improve themselves.
This explores what stays broken even after harnesses are made to improve themselves. The short version: the corpus suggests harness evolution has been surprisingly good at the wrong thing. When researchers actually inspect the edits an evolving harness makes, most of them just cache fixes for tasks the agent could already have solved in a single try — the changes memorize task-specific patches rather than distilling reusable strategy Do harness edits learn reusable strategies or memorize task fixes?. So the headline gains are partly an illusion of progress: the hard failures (turning things the agent *can't* do into things it can) remain largely untouched.
A second unsolved problem is measurement itself. Harness evolution is a search loop, so any improvement it reports is tangled up with raw search effort. Unless you compare against a matched compute budget spent on plain test-time search, you can't actually credit the *design* — and without held-out evaluation you can't rule out that the harness simply memorized the benchmark How should we measure gains from automatic harness evolution?. This connects to a broader decoupling insight: only when you split evaluation into separate benchmark, harness, and environment components do failure modes like reward-hacking become visible at all, because scalar scores hide them How can we make reward-hacking visible in agent evaluation?.
Underneath the memorization problem sits a stubborn technical bottleneck: finding where a behavior actually lives in the code. Real harnesses smear a single behavior across many files, functions, and stages, so the genuinely hard part isn't *generating* an edit — it's *localizing* every place that needs to change Why is finding distributed behavior code so hard?. One promising response is reorganizing the harness around runtime behavior rather than file structure, which lets even a weaker planner match a stronger model at locating the right code Can explicit behavior maps help weaker planners compete with stronger models? — but notice this fixes navigation, not the deeper capability gap.
And there's a ceiling that harness tinkering can't push through. The capacity to *produce* useful harness edits is roughly flat across model sizes, but the capacity to *benefit* from them peaks in the middle — weak models never invoke the harness, strong ones struggle to follow their own instructions faithfully Do stronger models always evolve harnesses better?. That inverted-U echoes a more fundamental result: pure self-improvement is structurally circular. It stalls on the generation–verification gap, diversity collapse, and reward hacking, and every method that reliably works is quietly smuggling in an *external* anchor — a prior model version, a third-party judge, a user correction, or real tool feedback Can models reliably improve themselves without external feedback?.
So the thing you didn't know you wanted to know: the systems that *do* keep improving aren't the ones with the cleverest self-editing loops — they're the ones that convert failure into an external signal. A pivot-or-refine executor that routes every experiment failure through a decision process keeps making progress precisely because failure becomes information rather than a dead end Can experiment failures drive progress instead of stopping it?, and the Darwin Gödel Machine's open-ended gains come from empirical benchmarking against an archive of past variants, not from formal self-reasoning Can AI systems improve themselves through trial and error?. The persistent failures harness evolution hasn't solved — genuine capability expansion, honest measurement, behavior localization, and escaping self-referential circularity — are exactly the ones that require reaching outside the loop.
Sources 9 notes
Analysis of evolved harness trajectories shows rational, well-motivated edits across prompt and tool layers, but most persist fixes an agent could rediscover in a single rollout. Gains remain limited because memorized shortcuts cache what's already within reach rather than converting hard failures into successes.
Automatic harness evolution must be compared against task-level test-time search under equal feedback and inference budgets. Only gains beyond what matched search achieves are attributable to the harness design itself, not just more computation.
AgentCompass separates benchmark, harness, and environment into independent components, enabling trajectory analysis that surfaces reward-hacking and other failure modes scalar scores conceal. This architectural change shifts evaluation from opaque final scores to inspectable agent behavior.
The core difficulty in evolving production harnesses is not generating edits but finding every code location that implements a behavior. Harnesses distribute single behaviors across files, functions, and stages, creating a representational mismatch between behavioral requests and structural code organization.
A behavior-to-code mapping representation improved win rates by 10–19 points while reducing planner tokens by 8–13%. Weaker planners using this mapping matched stronger models' code localization across all precision and recall metrics.
Show all 9 sources
Model capability to produce useful harness edits stays constant across tiers, but capacity to actually benefit from those edits follows an inverted U-shape, peaking in mid-tier models. Weak models fail to invoke harnesses; strong models struggle with faithful instruction-following.
Pure self-improvement stalls due to the generation-verification gap, diversity collapse, and reward hacking. Reliable improvement methods succeed by smuggling in external anchors: past model versions, third-party judges, user corrections, or tool feedback.
AutoResearchClaw's pivot-or-refine loop routes every failure through a decision process, making failure inform the next attempt rather than stop execution. Component ablation shows this mechanism drives completion and is distinct from reasoning or verification.
DGM replaces formal proofs with empirical benchmarking and maintains an evolutionary archive of agent variants, achieving 2.5× improvement on SWE-bench and 2.2× on Polyglot by discovering capabilities like better code editing and context management.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Rethinking the Evaluation of Harness Evolution for Agents
- Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
- Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
- DarwinX: Evolving Agent Harnesses Through Natural Selection
- HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
- Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution
- Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops
- RRSI: Regularized Recursive Self-Improvement of Agent Harnesses