Can explicit behavior maps help weaker planners compete with stronger models?
Explores whether organizing harness repositories around runtime behavior—rather than relying on model inference—can narrow the capability gap between weaker and stronger planning models, and whether this reduces computational overhead.
If Why is finding distributed behavior code so hard?, the fix is to build the missing map explicitly rather than make an agent recover it every time. Harness Handbook is a behavior-centric representation: it links behavioral descriptions to the distributed source that implements them, constructs that mapping directly from source, and supports editing through behavior-guided progressive disclosure plus automatic resynchronization when the code changes.
The results argue that the mapping — not raw model strength — was doing the work. With the handbook, overall win rates rise 10.0 and 18.9 points on Codex and Terminus-2 while planner token use falls 12.7% and 8.6%: better plans for less compute, because the planner no longer burns tokens rediscovering where behavior lives. The sharper evidence is substitutive: with the handbook a weaker planner matches the implementation-site localization of substantially stronger models, improving every file- and symbol-level Recall/Precision/F1 comparison against two independent reference plans, and the gains persist across request types and difficulty levels.
This is a concrete instance of a recurring pattern in the harness literature — capability relocating out of the model and into the scaffolding around it. Like Can externalized bookkeeping let smaller search agents beat larger ones?, the move is to externalize a structure (here, the behavior→code map) so the model spends its budget on judgment rather than bookkeeping. It also complements What happens to code that agents create and then share?: a durable, resynchronized representation is exactly the kind of persistent artifact that compounds.
Inquiring lines that read this note 46
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How does harness optimization generalize across different model architectures and domains?- Can harness updates benefit agents equally across all model sizes?
- How should harness scaffolding be treated as a first-class object?
- What makes harnesses more tangled than other types of agent code?
- Why does the harness layer accumulate distributed behaviors over time?
- Why do mid-tier models benefit more from memorized harness shortcuts?
- Can harness evolution be redirected toward distilling transferable procedures instead?
- What persistent failures remain unsolved despite harness evolution efforts?
- Why do evolved harness edits mostly memorize rather than generalize?
- What feedback signals matter most during harness evolution search?
- What makes behavior localization the bottleneck in agent harness evolution?
- How do different harness designs produce different agent behaviors from the same model?
- What makes a harness a first-class object rather than invisible scaffolding?
- Can harness evolution be redirected from memorization toward strategy distillation?
- How much of harness-evolution gain comes from matched test-time search budgets?
- Why does harness benefit capacity peak at mid-tier models, not frontier scale?
- Can runtime behavior mapping help localize harness deficiencies?
- How do evolved harness edits generalize across different benchmark domains?
- Can harness edits distill reusable strategies or mostly memorize task-specific fixes?
- How much realized agent capability comes from the harness versus the model?
- How do prompt optimization and code harnesses compare for capability transfer?
- What makes a harness low-friction for model strategy?
- Can weaker models match stronger ones by reorganizing harness-side components?
- Can mid-tier models benefit more from harness improvements than frontier models?
- Does harness scaling represent a fundamentally different path than model scaling?
- How much does harness design contribute to reported model capability scores?
- How does harness structure affect planner token efficiency compared to model size?
- What role does effective feedback compute play in agent harness scaling?
- Does harness optimization generalize across different benchmarks and agent architectures?
- Can weaker models benefit equally from harness updates as stronger ones?
- Can harness evolution gains be distinguished from test-time search improvements on matched budgets?
- How can we reorganize repositories to make behaviors easier to locate?
- Can weaker planners match stronger models if behavior is reorganized?
- Are durable shared code artifacts better than per-task harness patches?
- What makes durable code artifacts more valuable than per-task harness patches?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Why is finding distributed behavior code so hard?
When developers need to modify agent harnesses, they struggle to locate all the code implementing a target behavior because behaviors are scattered across files and stages while requests describe what to do, not where to look.
the problem this representation solves
-
Can externalized bookkeeping let smaller search agents beat larger ones?
Does offloading routine record-keeping to an environment harness free RL policies to focus on semantic search decisions, and can this approach outperform larger searchers with fewer parameters?
same externalize-structure-so-the-model-does-judgment pattern
-
What happens to code that agents create and then share?
Agent-authored code artifacts that persist across tasks and multiple agents remain poorly understood. The open questions cluster around what should be retained versus discarded, and how shared state stays consistent when multiple agents collaborate.
a persistent, compounding harness artifact
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
- An Empirical Study of Harness Design for Coding Agents
- From Model Scaling to System Scaling: Scaling the Harness in Agentic AI
- Scaling Laws for Agent Harnesses via Effective Feedback Compute
- Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
- StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
- Rethinking the Evaluation of Harness Evolution for Agents
- Code as Agent Harness
Original note title
reorganizing a harness repository around runtime behavior lets a weaker planner match a stronger model's code localization