Every time an AI agent breaks, the easiest fix gets sprinkled wherever it fits — and those patches compound into chaos.
Why does the harness layer accumulate distributed behaviors over time?
This explores why, as an agent's scaffolding (its prompts, tools, and orchestration code — the 'harness') evolves, single behaviors end up smeared across many files and stages rather than living in one clean place.
This explores why the harness — the scaffolding of prompts, tools, and orchestration code wrapped around a model — ends up with single behaviors scattered across many locations as it grows. The short answer from the corpus: harness evolution is an accretion process, and accretion favors local patches over global reorganization. When an agent hits a failure, the cheapest fix is to graft a small correction onto whatever stage was running — a prompt tweak here, a tool guard there. Over many rollouts these grafts pile up, and because a single 'behavior' (say, how the agent decides when to stop) touches the planner, a tool wrapper, and a post-processing step, the fix for it gets deposited in all three. Why is finding distributed behavior code so hard? names this directly: harnesses distribute one behavior across files, functions, and stages, creating a mismatch between how we *ask* for a behavior and how the code is actually organized.
The reason this keeps happening rather than self-correcting is that the edits are shallow by nature. Do harness edits learn reusable strategies or memorize task fixes? found that most evolved edits just cache a shortcut for an already-solvable task instead of distilling a reusable strategy — so the harness collects a sediment of narrow patches, each tied to a specific failure it once saw. There's no pressure toward consolidation because consolidating would mean *finding every place a behavior lives first*, which is exactly the hard part. That's why How should we measure gains from automatic harness evolution? warns that apparent gains can be memorization in disguise: the harness is quietly hoarding task-specific fixes, and only a held-out test reveals whether anything general was learned.
There's also a structural reason baked into how self-improving agents work. Do self-improving agents really split into two distinct loops? describes two loops: a slow one that rewrites model weights and a fast one that rewrites prompts, memory, and tools. Almost all recent progress runs through the fast loop precisely because scaffold edits are cheap and reversible — but 'cheap and reversible' is also what lets them accumulate unchecked. Weight updates force integration; scaffold edits don't. So behaviors that a weight-based learner might fold into a single representation instead stay externalized and spread out across the harness.
The corpus also points at the fix, which tells you something about the cause. Can explicit behavior maps help weaker planners compete with stronger models? shows that giving the harness an explicit behavior-to-code map — a representation that says 'this behavior is implemented in these places' — raised win rates and even let weaker planners match stronger ones at finding the right code. The fact that this helps so much confirms that the distribution isn't inevitable; it's what you get when nothing in the loop maintains a map of where behaviors live. Compare Can agents learn new skills without forgetting old ones?, where VOYAGER stores skills as discrete, indexed, composable units: when the substrate is designed for retrieval and reuse, accumulation produces a *library* instead of a tangle. The harness accumulates distributed behaviors because, unlike a skill library, it has no native unit of behavior — so growth means scattering, not filing.
The quietly interesting part: this isn't a bug you patch once. Because the capacity to *benefit* from harness edits peaks at mid-tier models and doesn't scale cleanly (Do stronger models always evolve harnesses better?), even a very strong model doesn't automatically clean up its own scaffolding. Distribution of behavior is the default state of any system that improves by patching itself faster than it reorganizes itself.
Sources 7 notes
The core difficulty in evolving production harnesses is not generating edits but finding every code location that implements a behavior. Harnesses distribute single behaviors across files, functions, and stages, creating a representational mismatch between behavioral requests and structural code organization.
Analysis of evolved harness trajectories shows rational, well-motivated edits across prompt and tool layers, but most persist fixes an agent could rediscover in a single rollout. Gains remain limited because memorized shortcuts cache what's already within reach rather than converting hard failures into successes.
Automatic harness evolution must be compared against task-level test-time search under equal feedback and inference budgets. Only gains beyond what matched search achieves are attributable to the harness design itself, not just more computation.
A survey framework organizes self-improving agents into two update mechanisms: slow parametric loops updating foundation model weights, and fast non-parametric loops updating prompts, memory, and tools. Recent progress concentrates in the fast loop because scaffold updates are cheaper and reversible than weight updates.
A behavior-to-code mapping representation improved win rates by 10–19 points while reducing planner tokens by 8–13%. Weaker planners using this mapping matched stronger models' code localization across all precision and recall metrics.
Show all 7 sources
VOYAGER demonstrates that storing executable skills in an embedding-indexed library and composing complex skills from simpler ones allows agents to learn continuously while avoiding the forgetting that occurs with weight-update-based methods. Environmental feedback refines skills while an automatic curriculum drives continual exploration.
Model capability to produce useful harness edits stays constant across tiers, but capacity to actually benefit from those edits follows an inverted U-shape, peaking in mid-tier models. Weak models fail to invoke harnesses; strong models struggle with faithful instruction-following.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Rethinking the Evaluation of Harness Evolution for Agents
- Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
- Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
- HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
- Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution
- DarwinX: Evolving Agent Harnesses Through Natural Selection
- StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
- ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement