If a fix is tuned on one AI model, how much of its benefit survives when you move it to a different model?
How much do cross-model improvements compare to the original performance gains on the training benchmark?
This explores what happens when an improvement tuned on one model or one benchmark is carried over to other models or unseen tasks: does the gain hold its size, shrink, or vanish?
This explores whether improvements built on one model or one benchmark keep their size when moved somewhere new, or whether most of the gain stays home. The short honest answer is that the collection shows transfer clearly happens, but it rarely puts the original gain and the transferred gain side by side as a ratio. Read the numbers below as a pattern, not a precise exchange rate.
The clearest case is harness work, meaning changes to the code and scaffolding around a model rather than its weights. One execution-harness runbook built for Terminal-Bench 2.1 was applied unchanged to newer models: it reached 95.3% on GPT-5.6 and lifted DeepSeek-V4 Flash by 5.4 points Can execution harnesses lift model performance without retuning weights?. Model-to-model gains like these tend to be real but uneven. A frontier model that is already strong lands near the ceiling, while a smaller model gets a modest bump. A related result shows how big the lift can be when the receiving model is weak: harnesses built by a stronger model nearly doubled a weaker model's Theory-of-Mind scores. The gain came from moving shaky reasoning into deterministic code, not from asking the model to think harder Can a stronger model lift a weaker one at test time without retraining?. That points to a useful rule of thumb: the improvements that transfer best are the ones that take work away from the model instead of tuning it.
Transfer across benchmarks depends on how the improvement was found. AIDE2's gains held up on four held-out benchmarks, including weather forecasting, which sits well outside the tasks it was tuned on Do AIDE2's improvements transfer to unseen tasks?. ModularRSI builds this in from the start. It improves harness components on data kept separate from the benchmark and pools evidence across tasks before changing anything, specifically to separate reusable mechanisms from benchmark-specific tricks Can harness modules improve separately from benchmark data?. Feedback quality matters too. A small 9B editor trained on whether its patches actually worked gave a steady 9.3-point lift across three tasks. Prompted frontier models gave unstable or smaller gains because they optimized for patches that looked plausible, not ones that were verified Does training editors on real outcomes beat prompting larger models?.
A result from a different field gives a useful benchmark for how much of a gain typically survives a handoff. When an LLM got better at individual questions, the humans using it captured only about half of that improvement Why does assisted accuracy capture only half the LLM gain?. Something similar happens in training. Higher-quality teacher data can make a student model worse when it goes beyond what the student can absorb Does teacher-refined data always improve student model performance?. Both point the same way: a gain measured where it was made is a ceiling, not a promise. How much survives depends on whether the receiving side can use it.
The takeaway you may not have expected: the most portable improvements tend to be the least model-specific ones. They include deterministic code, checks that can be verified, and components tuned on data separate from the benchmark. Tuning toward a single benchmark score is what makes gains stay home. If you want exact figures comparing the original gain with the transferred gain, this collection doesn't provide them yet.
Sources 7 notes
StateM improves Terminal-Bench 2.1 accuracy across multiple models by optimizing execution systems around frozen weights. The same runbook transfers to newer models without modification, achieving 95.3% on GPT-5.6 and lifting DeepSeek-V4 Flash by 5.4 points.
A stronger model built inference-time harnesses that nearly doubled weaker model performance on Theory-of-Mind benchmarks without retraining, primarily by moving unstable reasoning into deterministic code and task-specific routing rather than encouraging extended reasoning.
The paper reports that AIDE2's improvements transfer to four held-out benchmarks spanning machine learning, algorithm engineering, and physics-based weather forecasting—the last being outside the selection task distribution. This demonstrates transferable gains beyond overfitting to the selection set.
ModularRSI evolves harness modules independently using contrastive trajectories on benchmark-disjoint data, showing consistent gains across unseen tasks and domains. The approach isolates mechanism-level improvements from task-specific adaptation by aggregating evidence across tasks before updating components.
A 9B model trained with reinforcement learning on patch success raised a frozen agent's performance by 9.3 points across three tasks, while prompted frontier models produced unstable or lower gains. The difference stems from feedback: trained editors rerun patches to verify impact, while prompted models optimize only for plausibility.
Show all 7 sources
A 535-participant study found that when LLM accuracy improved on individual items, assisted participants captured roughly half that gain—falling below what the better-performing component could have provided alone. This shows complementarity creates potential but does not guarantee synergy.
Teacher-refined data degrades performance when it exceeds the student's learning frontier, even if objectively higher quality. Students should filter refinements using their own statistical profile to retain only compatible improvements.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Sharpening Tax in Post-Training
- DarwinX: Evolving Agent Harnesses Through Natural Selection
- Scaling Laws for Agent Harnesses via Effective Feedback Compute
- Available but Unclaimed: An Empirical Study of Human-AI Synergy
- Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
- AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
- HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?