Two AIs with identical weights can score wildly differently on a benchmark — just by changing the code wrapped around them.
How much does harness design separate from model intelligence affect benchmark scores?
This explores how much of a model's benchmark score comes from the scaffolding around it (prompts, code, routing, tools, execution loops), as opposed to the model's own built-in capability.
This explores how much of a benchmark score really comes from the model, and how much from the 'harness': the code, prompts, tool routing and execution loop wrapped around it. The corpus says it's a lot more than most leaderboards suggest. On Terminal-Bench 2.1, improving only the execution system around frozen models raised accuracy for several of them, and the same runbook carried over to newer models with no changes Can execution harnesses lift model performance without retuning weights?. In a more striking case, a stronger model wrote inference-time harnesses that nearly doubled a weaker model's scores on Theory-of-Mind tests, without any retraining Can a stronger model lift a weaker one at test time without retraining?. So the same weights can produce very different numbers depending on what surrounds them.
The interesting part is *why* harnesses help. They don't seem to make the model think harder. They take shaky reasoning away from it. In the Theory-of-Mind work, the gains came mostly from moving unstable reasoning into deterministic code and sending each kind of task down its own path. That fits the finding that splitting a 'decomposer' (which plans) from a 'solver' (which executes) beats one model doing both, and that planning skill transfers across domains while solving skill doesn't Does separating planning from execution improve reasoning accuracy?. A similar effect shows up in code work: giving a weaker planner a map from runtime behaviors to the code responsible for them let it match stronger models at finding the right code, while using fewer tokens Can explicit behavior maps help weaker planners compete with stronger models?. In each case the harness supplies structure that would otherwise have to come from raw intelligence.
The relationship isn't a simple 'better harness, better score for everyone,' though. One study found that models at every tier are about equally good at *suggesting* useful harness changes. But the benefit models get from those changes follows an inverted U. Weak models fail to use the harness at all. Strong models drift from its instructions. Mid-tier models gain the most Do stronger models always evolve harnesses better?. Safety harnesses show the same fit problem: a policy tuned for one model blocks too much on another, so they have to be tuned per deployment Should safety harnesses be customized for each deployment?.
This makes measurement hard. If a harness was evolved against a benchmark, part of its gain may be extra compute or overfitting to that benchmark. The corpus argues harness gains only count once they beat test-time search given the same budget How should we measure gains from automatic harness evolution?. It also argues for evolving harness parts on data kept separate from the benchmark, so reusable improvements can be told apart from task-specific tuning Can harness modules improve separately from benchmark data?. The practical result: a score that doesn't record the setup it came from is hard to compare with anything Can benchmark scores be trusted without knowing their origin?.
The takeaway you might not have expected is that the harness isn't just a measurement nuisance. It may be where AI first starts improving itself. Lilian Weng argues that recursive self-improvement in the near term will probably come from models improving their own prompts, harness code and optimizers, not from rewriting their own weights Does recursive self-improvement start with harness engineering?. If so, the gap between 'model intelligence' and 'harness design' is likely to blur further over time. The corpus doesn't give one percentage split, and the evidence suggests no single number exists, because the split depends on model tier, task and budget.
Sources 10 notes
StateM improves Terminal-Bench 2.1 accuracy across multiple models by optimizing execution systems around frozen weights. The same runbook transfers to newer models without modification, achieving 95.3% on GPT-5.6 and lifting DeepSeek-V4 Flash by 5.4 points.
A stronger model built inference-time harnesses that nearly doubled weaker model performance on Theory-of-Mind benchmarks without retraining, primarily by moving unstable reasoning into deterministic code and task-specific routing rather than encouraging extended reasoning.
Modular architectures with separate decomposer and solver models outperform monolithic LLMs, with decomposition ability transferring across domains while solving ability does not. The separation prevents planning-execution interference and produces more generalizable skills.
A behavior-to-code mapping representation improved win rates by 10–19 points while reducing planner tokens by 8–13%. Weaker planners using this mapping matched stronger models' code localization across all precision and recall metrics.
Model capability to produce useful harness edits stays constant across tiers, but capacity to actually benefit from those edits follows an inverted U-shape, peaking in mid-tier models. Weak models fail to invoke harnesses; strong models struggle with faithful instruction-following.
Show all 10 sources
A harness strict enough for one model over-blocks another, while policies general enough to transfer across domains miss application-specific safety relations. Domain semantics and model characteristics jointly determine which harness is effective.
Automatic harness evolution must be compared against task-level test-time search under equal feedback and inference budgets. Only gains beyond what matched search achieves are attributable to the harness design itself, not just more computation.
ModularRSI evolves harness modules independently using contrastive trajectories on benchmark-disjoint data, showing consistent gains across unseen tasks and domains. The approach isolates mechanism-level improvements from task-specific adaptation by aggregating evidence across tasks before updating components.
Benchmark Radar catalogs AI evaluations while keeping source identities and citations attached, enabling readers to trace scores back to their original settings. The system demonstrates that scores without provenance cannot reliably support model comparisons.
Weng argues RSI's initial path moves through optimizing deployment harnesses—instruction prompts to harness code to optimizer code—rather than models directly rewriting weights. This staged progression mirrors how prompt engineering gave way to instruction tuning while interface needs persisted.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
- Rethinking the Evaluation of Harness Evolution for Agents
- HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
- Sharpening Tax in Post-Training
- StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
- Harness Engineering for Self-Improvement
- Scaling Laws for Agent Harnesses via Effective Feedback Compute
- Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution