SYNTHESIS NOTE
Topics›Agent Harness›this note

Do stronger models always evolve harnesses better?

We explore whether base model capability predicts both the ability to write useful harness updates and the ability to benefit from them. The answer reshapes how we should allocate capability in self-evolving agent systems.

Synthesis note · 2026-06-03 · sourced from Agent Harness

Self-evolving agents edit an external harness — prompts, skills, memories, tools — from execution evidence, without touching model parameters. The natural assumption is that stronger base models do this better on both ends. This paper disentangles two distinct capabilities and finds neither follows that assumption.

Harness-updating — producing persistent edits that lead to gains — is flat in base capability. Models across capability tiers produce updates yielding surprisingly similar gains; even a Qwen3.5-9B evolver induces gains comparable to Claude Opus 4.6. Writing a good skill or memory is apparently not bottlenecked by raw model strength.

Harness-benefit — actually improving when handed an updated harness — is non-monotonic. Weak-tier models gain little, mid-tier models benefit most, and strong-tier models benefit less than mid-tier. Two failure modes explain the weak end: failing to activate the relevant harness artifact, and failing to follow it faithfully once activated.

The practical inversion is sharp: invest capability budget in the agent that uses the harness, not the evolver that writes it — and target agent training at harness invocation and long-horizon instruction-following rather than at generating cleverer updates. This complicates the "let a frontier model improve everything" intuition and connects to Why do better reasoning models ignore instructions?: strong models may benefit less precisely because the bottleneck is faithful instruction-following, which scaling erodes.

Inquiring lines that read this note 96

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How does harness optimization generalize across different model architectures and domains? What fundamental constraints limit how effectively agents can improve themselves? Does abstract user knowledge outperform concrete interaction history in personalization? How should agent systems validate and persist generated code artifacts? What execution architectures enable agents to most effectively use tools? Why do standard benchmarks fail to predict agent deployment success? How do agent-learned skills transfer and improve across different tasks? How do standardized protocols improve multi-agent coordination and reliability? What reasoning architectures enable models to solve complex problems efficiently? Can harness architecture and protocols provide agent reliability without model scaling? When do multi-agent systems outperform single frontier models? What capability trade-offs arise from domain specialization through fine-tuning? Why do agents falsely report success on failed tasks? What should agent evaluation prioritize to reveal reliable behavior? Do multi-agent systems introduce security vulnerabilities that single-agent architectures avoid? How can humans maintain meaningful oversight as AI systems become increasingly autonomous and complex? How do neighboring agents influence whether others cooperate or collude? What emerges when safety-aligned models attempt to role-play deceptive personas? How do training data properties determine the emergence of internal misalignment? How do we enforce security boundaries in evaluation environments? How do capability benchmark scores systematically misrepresent true model abilities? How can evolutionary algorithms maintain diversity during solution search? How can infrastructure records verify actual agent behavior? Can prompt-based context override biases that were embedded during pretraining? How do multi-agent LLM systems fail distinctly compared to single agents? Should agents decouple planning from perception grounding for better performance?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
16 direct connections · 76 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

the capacity to produce useful harness updates is flat across model tiers but the capacity to benefit from them peaks at mid-tier