When you make an AI's tools and prompts better, does that improvement ever get baked into the AI itself?
Do harness improvements eventually internalize into core model behavior over time?
This explores whether the gains people get from improving an AI's harness (the prompts, tools, memory handling and control code wrapped around a model) eventually get absorbed into the model itself, so the harness stops mattering, or whether harness and model stay separate layers.
This explores whether improvements to the scaffolding around a model (prompts, tools, context handling, control flow) eventually get absorbed into the model's weights, so the scaffolding becomes unnecessary. No study in this collection tracks that absorption directly over time. What it does have points to a split answer: some harness improvements look like temporary workarounds that training will likely take over, while others look permanent because they do jobs that model weights don't do well.
The clearest argument that absorption happens comes from Weng's account of recursive self-improvement Does recursive self-improvement start with harness engineering?. It points to a precedent: prompt engineering tricks were later built into models through instruction tuning. Weng adds a caveat, though. Even after that happened, models still needed an interface around them. On that view the harness is a testing ground. Behaviors are tried there first, some later move into the model, and a working layer stays behind. A useful test for which harness edits are likely to move is in Do harness edits learn reusable strategies or memorize task fixes?. Most evolved edits just store fixes the agent could have found again on its own in a single attempt. Fixes like that are exactly what a stronger or better-trained model should stop needing. Related work shows that harness self-improvement easily overfits the tasks it was tuned on unless it is constrained to favor reusable mechanisms Does harness self-improvement memorize tasks instead of learning broadly? Can harness modules improve separately from benchmark data?.
Other harness gains look like they won't be absorbed. Auto-research loops run across many environments found four mechanisms that cut token traffic by roughly 45–49%: context compaction, observation handling, delegated reading and action execution. The authors describe these gains as independent of model improvements Can agent harnesses be automatically optimized across many environments?. Those are mostly engineering decisions about what information reaches the model, not behaviors the model could learn. In the same vein, a stronger model nearly doubled a weaker model's Theory-of-Mind scores. Most of that gain came from moving unstable reasoning out of the model and into deterministic code Can a stronger model lift a weaker one at test time without retraining?. That is the opposite of internalizing: work is deliberately taken away from the weights because code does it more reliably.
The surprising finding is that pushing harness gains into the weights through more training doesn't reliably work right now. In RSIGym's joint trials, harness revisions consistently raised scores, while heavier fine-tuning lowered them in eight of ten comparisons Do harness fixes or heavier training drive frontier model gains?. Frozen models gained an average of 17 points from harness evolution alone Can frozen models improve by evolving their harnesses?. One paper finds that how much a model benefits from harness edits peaks at mid-tier models Do stronger models always evolve harnesses better?. Weak models fail to use the harness, and the strongest models have trouble following its instructions faithfully. That fits partial absorption: as models get stronger, some of the scaffolding starts to get in the way.
The idea worth keeping is that 'harness versus model' may be the wrong way to frame it. Macaron-V1 argues that adaptation belongs to the whole versioned loop of model, harness and external contract, not to any single model snapshot. Weight updates in that design only happen after passing audits Where does model adaptation actually happen?. If that is right, harness improvements don't disappear into the model over time. They move between layers. Workarounds for specific tasks drift into the weights, and context management, deterministic code and checks against outside requirements stay in the harness.
Sources 10 notes
Weng argues RSI's initial path moves through optimizing deployment harnesses—instruction prompts to harness code to optimizer code—rather than models directly rewriting weights. This staged progression mirrors how prompt engineering gave way to instruction tuning while interface needs persisted.
Analysis of evolved harness trajectories shows rational, well-motivated edits across prompt and tool layers, but most persist fixes an agent could rediscover in a single rollout. Gains remain limited because memorized shortcuts cache what's already within reach rather than converting hard failures into successes.
Recursive edits to agent scaffolding can memorize evolve-set tasks, causing in-distribution gains to shrink out-of-distribution. RRSI addresses this by constraining the proposal and selection stages to favor reusable mechanisms over benchmark-specific ones.
ModularRSI evolves harness modules independently using contrastive trajectories on benchmark-disjoint data, showing consistent gains across unseen tasks and domains. The approach isolates mechanism-level improvements from task-specific adaptation by aggregating evidence across tasks before updating components.
Cross-environment harness optimization yielded four mechanisms (action execution, context compaction, observation handling, delegated reading) that reduced token traffic by 44.7–49.0% while maintaining comparable performance on a 51-task benchmark, suggesting harness-level gains are orthogonal to model improvements.
Show all 10 sources
A stronger model built inference-time harnesses that nearly doubled weaker model performance on Theory-of-Mind benchmarks without retraining, primarily by moving unstable reasoning into deterministic code and task-specific routing rather than encouraging extended reasoning.
Across six frontier models in RSIGym's Joint track, harness fixes consistently improved performance while additional fine-tuning lowered scores in eight of ten comparisons. Claude agents ranked highest by spending more time diagnosing and revising harnesses rather than scaling training.
DarwinX achieves average 17-point gains across benchmarks by evolving harness variants (prompts, tools, skills, control flow) under a preserve-and-extend contract while keeping the model frozen. Key evidence includes Terminal-Bench 2.1 rising to 84.7% and WebArena-Infinity reaching 93.0% audit-clean pass@1.
Model capability to produce useful harness edits stays constant across tiers, but capacity to actually benefit from those edits follows an inverted U-shape, peaking in mid-tier models. Weak models fail to invoke harnesses; strong models struggle with faithful instruction-following.
Macaron-V1 argues adaptation is a property of the recursive cycle linking model, harness, and external contract—not individual model snapshots. Weight updates are gated by audit and evaluation against an external contract, making the loop the unit of improvement and release.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
- HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
- Rethinking the Evaluation of Harness Evolution for Agents
- DarwinX: Evolving Agent Harnesses Through Natural Selection
- RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
- Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution
- ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement
- Sharpening Tax in Post-Training