Do stronger models always evolve harnesses better?
We explore whether base model capability predicts both the ability to write useful harness updates and the ability to benefit from them. The answer reshapes how we should allocate capability in self-evolving agent systems.
Self-evolving agents edit an external harness — prompts, skills, memories, tools — from execution evidence, without touching model parameters. The natural assumption is that stronger base models do this better on both ends. This paper disentangles two distinct capabilities and finds neither follows that assumption.
Harness-updating — producing persistent edits that lead to gains — is flat in base capability. Models across capability tiers produce updates yielding surprisingly similar gains; even a Qwen3.5-9B evolver induces gains comparable to Claude Opus 4.6. Writing a good skill or memory is apparently not bottlenecked by raw model strength.
Harness-benefit — actually improving when handed an updated harness — is non-monotonic. Weak-tier models gain little, mid-tier models benefit most, and strong-tier models benefit less than mid-tier. Two failure modes explain the weak end: failing to activate the relevant harness artifact, and failing to follow it faithfully once activated.
The practical inversion is sharp: invest capability budget in the agent that uses the harness, not the evolver that writes it — and target agent training at harness invocation and long-horizon instruction-following rather than at generating cleverer updates. This complicates the "let a frontier model improve everything" intuition and connects to Why do better reasoning models ignore instructions?: strong models may benefit less precisely because the bottleneck is faithful instruction-following, which scaling erodes.
Inquiring lines that read this note 96
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How does harness optimization generalize across different model architectures and domains?- Can mid-tier models benefit more from self-generated harness updates than others?
- Can smaller models produce skill updates as useful as frontier model updates?
- What happens when different harnesses project the same model?
- Does harness benefit depend on which model tier you use?
- What makes skills worth externalizing into a persistent harness?
- What cognitive burdens should move from model parameters into harness infrastructure?
- What causes weak models to fail at activating harness artifacts?
- How should we allocate model budget between evolvers and harness users?
- Can harness updates benefit agents equally across all model sizes?
- How should harness scaffolding be treated as a first-class object?
- Why does the harness layer accumulate distributed behaviors over time?
- Why do mid-tier models benefit more from memorized harness shortcuts?
- Can harness evolution be redirected toward distilling transferable procedures instead?
- What persistent failures remain unsolved despite harness evolution efforts?
- Why do evolved harness edits mostly memorize rather than generalize?
- What feedback signals matter most during harness evolution search?
- What makes behavior localization the bottleneck in agent harness evolution?
- Why do persistent, resynchronized artifacts compound harness capability gains?
- Do evolved harness edits capture reusable strategies or task-specific memorization?
- How do different harness designs produce different agent behaviors from the same model?
- What makes a harness a first-class object rather than invisible scaffolding?
- Why do mid-tier models benefit most from memorized harness fixes?
- Can harness evolution be redirected from memorization toward strategy distillation?
- How much of harness-evolution gain comes from matched test-time search budgets?
- Why do evolved harnesses often fail to generalize beyond their training tasks?
- Why does harness benefit capacity peak at mid-tier models, not frontier scale?
- How does editing the harness layer differ from updating model weights?
- Which foundation model tiers most benefit from harness updates?
- Can runtime behavior mapping help localize harness deficiencies?
- How do evolved harness edits generalize across different benchmark domains?
- Can harness edits distill reusable strategies or mostly memorize task-specific fixes?
- How much realized agent capability comes from the harness versus the model?
- How do prompt optimization and code harnesses compare for capability transfer?
- What should an external contract for model improvement actually contain?
- How do agentic systems hide harness failures from benchmarks?
- What makes a harness low-friction for model strategy?
- Can weaker models match stronger ones by reorganizing harness-side components?
- Which domains see models exceed human harness design quality?
- Why do useful harness updates often disappear during model evolution?
- How much does executor choice change a harness's actual performance?
- Do models co-adapt their harnesses to specific executor strengths?
- Can mid-tier models benefit more from harness improvements than frontier models?
- Does harness scaling represent a fundamentally different path than model scaling?
- How much does harness design contribute to reported model capability scores?
- How does harness structure affect planner token efficiency compared to model size?
- What role does effective feedback compute play in agent harness scaling?
- Does harness optimization generalize across different benchmarks and agent architectures?
- Can weaker models benefit equally from harness updates as stronger ones?
- Can harness evolution gains be distinguished from test-time search improvements on matched budgets?
- Should we train the evolver or the executor when building self-improving agents?
- How can agents evolve their own skills without human input?
- Does removing static external utility break the formal guarantees of self-improvement loops?
- Why do self-improving agents concentrate progress in the fast non-parametric loop?
- How many acceptable rewrites can recursive self-improvement sustain before returns diminish?
- What makes an agent in an economic simulation self-evolving?
- Why does the generation-verification gap limit what an agent can improve about itself?
- Do evolutionary archives let agents improve themselves without formal proof?
- How do agent-created code artifacts become part of harness infrastructure?
- Are durable shared code artifacts better than per-task harness patches?
- Can disposable agent-authored code be distinguished from reusable infrastructure?
- What makes durable code artifacts more valuable than per-task harness patches?
- Can single-axis benchmarks measure across all three agent capability layers?
- Does a single benchmark score systematically misrepresent multi-axis agent capability?
- Which separable capability axes reveal when single benchmarks misrepresent deployment readiness?
- Can agent-authored skill libraries compound autonomy gains over time?
- What stops evolved agent behaviors from generalizing beyond specific tasks?
- Can skill repositories evolve toward execution-oriented refinement over time?
- How much does external context management transfer across similar capability agents?
- Do task-level outcomes provide sufficient supervision for harness evolution?
- How do parametric and non-parametric updates differ in agents?
- How should versioning and rollback govern the fast scaffold update loop?
- Which harness dimensions most directly predict agent system reliability?
- Can slower development eliminate the risk of failure in agentic systems?
- How does role allocation in multi-agent systems depend on model differentiation?
- Do single-agent systems outperform multi-agent coordination as model capabilities grow?
- How much does workflow architecture matter versus raw model capability?
- What role does pretraining play in distinguishing system capability from deployed behavior?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Where does agent reliability actually come from?
Exploring whether LLM agent performance depends on larger models or on thoughtful system design choices like memory, skills, and protocols that shift cognitive work outside the model.
sharpens it: which model should hold which role in a harness-evolution loop
-
Can a separate trained curator improve skill libraries better than frozen agents?
Explores whether decoupling skill curation from agent execution enables better long-term learning of what skills to keep, delete, or refine. Matters because manual curation doesn't scale and heuristic approaches lack feedback.
same updater/executor split; this note adds the capability-tier asymmetry
-
How can agent self-evolution be made safe and auditable?
As agents begin updating their own prompts and tools, how can we track these changes, measure their effects, and safely reverse problematic updates? This matters because untracked evolution leads to unmaintainable systems and makes regressions impossible to diagnose.
the protocol substrate this capability analysis runs on top of
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
- Rethinking the Evaluation of Harness Evolution for Agents
- HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
- Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
- RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
- Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution
- StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
- DarwinX: Evolving Agent Harnesses Through Natural Selection
Original note title
the capacity to produce useful harness updates is flat across model tiers but the capacity to benefit from them peaks at mid-tier