Do harness edits learn reusable strategies or memorize task fixes?
When meta-agents evolve harnesses iteratively, do the persisted edits encode transferable procedures that solve new problems, or do they mostly cache shortcuts for already-solvable tasks? This matters because it determines whether harness evolution genuinely expands capability.
When you open up the trajectories a harness-evolution meta-agent produces — what it modifies each iteration, which edits get kept or rolled back, which failure classes they target — the edits look rational and well-motivated across the prompt, middleware, and tool layers. Yet a stable core of hard tasks stays unsolved and the overall improvement remains limited relative to simple test-time discovery baselines. The diagnosis is sharp: most edits memorize fixes rather than distill strategies. Much of what gets written into the harness is information a competent agent could rediscover through exploration inside a single rollout, so persisting it saves time on tasks the agent could already solve but rarely converts a failure into a success.
This reframes what harness evolution is buying. The value is not "the harness learned to do new things" but "the harness caches shortcuts for things already within reach." That is why gains shrink once you control the search budget — see How should we measure gains from automatic harness evolution?.
It also sharpens a companion finding: since Do stronger models always evolve harnesses better?, the payoff concentrates exactly where memorized shortcuts substitute for capability the agent lacks. The open question is whether harness evolution can be redirected from memorization toward strategy distillation — edits that encode transferable procedures rather than per-task patches — which is what would actually move the stable core of hard tasks.
Inquiring lines that read this note 32
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What fundamental constraints limit how effectively agents can improve themselves?- What makes evolving the benchmark different from evolving the optimizer itself?
- Does AIDE2 archive rejected variants the way evolutionary approaches do for future reuse?
- Can harness updates benefit agents equally across all model sizes?
- How should harness scaffolding be treated as a first-class object?
- What makes harnesses more tangled than other types of agent code?
- Why does the harness layer accumulate distributed behaviors over time?
- Can harness evolution be redirected toward distilling transferable procedures instead?
- What persistent failures remain unsolved despite harness evolution efforts?
- Why do evolved harness edits mostly memorize rather than generalize?
- What feedback signals matter most during harness evolution search?
- What makes behavior localization the bottleneck in agent harness evolution?
- Do evolved harness edits capture reusable strategies or task-specific memorization?
- How do different harness designs produce different agent behaviors from the same model?
- What makes a harness a first-class object rather than invisible scaffolding?
- Why do mid-tier models benefit most from memorized harness fixes?
- Can harness evolution be redirected from memorization toward strategy distillation?
- How much of harness-evolution gain comes from matched test-time search budgets?
- Why do evolved harnesses often fail to generalize beyond their training tasks?
- How does editing the harness layer differ from updating model weights?
- Can runtime behavior mapping help localize harness deficiencies?
- How do evolved harness edits generalize across different benchmark domains?
- Can harness edits distill reusable strategies or mostly memorize task-specific fixes?
- Can harness edits trained on one batch transfer to new tasks?
- How do prompt optimization and code harnesses compare for capability transfer?
- What makes a harness low-friction for model strategy?
- Why do useful harness updates often disappear during model evolution?
- How much does executor choice change a harness's actual performance?
- Does harness optimization generalize across different benchmarks and agent architectures?
- Can harness evolution gains be distinguished from test-time search improvements on matched budgets?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
How should we measure gains from automatic harness evolution?
Harness evolution itself runs a search loop, so reported improvements might come from more search rather than better design. What's the right way to measure whether the harness itself actually improved?
explains why the measured gains stay small once budget is matched
-
Do stronger models always evolve harnesses better?
We explore whether base model capability predicts both the ability to write useful harness updates and the ability to benefit from them. The answer reshapes how we should allocate capability in self-evolving agent systems.
where memorized-fix payoff concentrates
-
What happens to code that agents create and then share?
Agent-authored code artifacts that persist across tasks and multiple agents remain poorly understood. The open questions cluster around what should be retained versus discarded, and how shared state stays consistent when multiple agents collaborate.
the contrast: durable reusable artifacts vs memorized per-task patches
-
Can prompt optimization accidentally teach judges to reward the wrong signals?
When prompts are persistently revised to improve a score, the optimization might find shortcuts that satisfy a judge's preferences without improving actual task performance. This matters because shortcuts embedded in reused instructions affect every downstream input, not just one interaction.
a persisted edit fitted to the evaluator and not to task instances: the mutation adopted a judge's preferred vocabulary while the second measure stayed flat; relayed, system undescribed, and not shown to be a task-specific fix
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Rethinking the Evaluation of Harness Evolution for Agents
- Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
- Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution
- HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
- DarwinX: Evolving Agent Harnesses Through Natural Selection
- Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
- ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement
- Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design
Original note title
evolved harness edits mostly memorize task-specific fixes rather than distilling reusable strategies