INQUIRING LINE

An AI that keeps rewriting its own toolkit might just be memorizing the test, not actually getting smarter — how would you tell?

Can harness edits overfit to training tasks while appearing to improve performance?

This explores whether automatically editing an agent's harness (the prompts, tools and scaffolding around a frozen model) can look like progress on the tasks it was tuned on while really just memorizing them, and what the corpus says about telling real improvement from overfitting.


This explores whether changes to an agent's harness (the prompts, tools and control code wrapped around a model) can raise scores on the tasks they were tuned against without making the agent better in general. The corpus says yes, and it's the usual outcome. When an agent repeatedly rewrites its own scaffolding against a fixed set of tasks, the edits tend to memorize those tasks. The gains on familiar problems shrink once the agent meets new ones Does harness self-improvement memorize tasks instead of learning broadly?. It's the same overfitting that happens in model training. The difference is that the 'weights' are text and code.

The edits don't look like overfitting. When researchers read through evolved harness changes, most of them were sensible, well-reasoned fixes to prompts and tools. But most of those fixes only saved the agent work it could have redone itself within a single attempt. The harness was caching answers to problems the agent could already solve, not teaching it to solve harder ones Do harness edits learn reusable strategies or memorize task fixes?. So reading the edits won't tell you whether they generalize. A reasonable-sounding change can still be a shortcut. A related result shows why: frontier models asked to patch a harness produce edits that look plausible, but the gains are unstable. A small model trained on whether its patches actually worked, which reruns them to check, does better Does training editors on real outcomes beat prompting larger models?. If nothing tests whether an edit really helps, plausibility is all the editor optimizes for.

The fixes the corpus describes look like the standard tools against overfitting, carried over to text edits. One approach changes how edits are proposed and chosen so that reusable mechanisms win over benchmark-specific tweaks Does harness self-improvement memorize tasks instead of learning broadly?. Another evolves each harness module separately on data kept apart from the benchmark. It compares successful and failed runs and pools evidence across many tasks before changing anything, which keeps task-specific adaptation from creeping in Can harness modules improve separately from benchmark data?. Work on skill rewriting adds a 'learning rate' that caps how much an edit can change, a held-out validation check, and a buffer that keeps rejected edits as negative examples. The skills that result are more stable and generalize better than when agents freely rewrite their own instructions Does constraining edits make skill learning more stable?.

None of this means harness work is fake progress. Some harness gains clearly do transfer. In one study, the same runbook carried over unchanged to newer models and still lifted their scores Can execution harnesses lift model performance without retuning weights?. Optimizing across many environments at once turned up general mechanisms, such as context compaction and delegated reading, that cut token use roughly in half Can agent harnesses be automatically optimized across many environments?. In a head-to-head trial, harness revisions beat extra fine-tuning Do harness fixes or heavier training drive frontier model gains?. The pattern across these results is that gains generalize when edits are tested on unseen tasks, other models or many environments. Gains measured only on the tasks being tuned against are suspect.

One harness result shows where the line sits. A stronger model built harnesses that nearly doubled a weaker model's scores on Theory-of-Mind benchmarks. It did this mostly by moving shaky reasoning into deterministic code and adding routing rules specific to each task Can a stronger model lift a weaker one at test time without retraining?. Whether you call that capability transfer or very good overfitting depends on whether the tasks you care about look like the benchmark. That's the question to ask of any reported harness gain.


Sources 9 notes

Does harness self-improvement memorize tasks instead of learning broadly?

Recursive edits to agent scaffolding can memorize evolve-set tasks, causing in-distribution gains to shrink out-of-distribution. RRSI addresses this by constraining the proposal and selection stages to favor reusable mechanisms over benchmark-specific ones.

Do harness edits learn reusable strategies or memorize task fixes?

Analysis of evolved harness trajectories shows rational, well-motivated edits across prompt and tool layers, but most persist fixes an agent could rediscover in a single rollout. Gains remain limited because memorized shortcuts cache what's already within reach rather than converting hard failures into successes.

Does training editors on real outcomes beat prompting larger models?

A 9B model trained with reinforcement learning on patch success raised a frozen agent's performance by 9.3 points across three tasks, while prompted frontier models produced unstable or lower gains. The difference stems from feedback: trained editors rerun patches to verify impact, while prompted models optimize only for plausibility.

Can harness modules improve separately from benchmark data?

ModularRSI evolves harness modules independently using contrastive trajectories on benchmark-disjoint data, showing consistent gains across unseen tasks and domains. The approach isolates mechanism-level improvements from task-specific adaptation by aggregating evidence across tasks before updating components.

Does constraining edits make skill learning more stable?

SkillOpt's ablations show that adding a textual learning-rate budget, held-out validation gate, and rejected-edit buffer (retaining failed edits as negative feedback) produces more stable and generalizable skill improvement than allowing agents to freely rewrite their own instructions.

Show all 9 sources
Can execution harnesses lift model performance without retuning weights?

StateM improves Terminal-Bench 2.1 accuracy across multiple models by optimizing execution systems around frozen weights. The same runbook transfers to newer models without modification, achieving 95.3% on GPT-5.6 and lifting DeepSeek-V4 Flash by 5.4 points.

Can agent harnesses be automatically optimized across many environments?

Cross-environment harness optimization yielded four mechanisms (action execution, context compaction, observation handling, delegated reading) that reduced token traffic by 44.7–49.0% while maintaining comparable performance on a 51-task benchmark, suggesting harness-level gains are orthogonal to model improvements.

Do harness fixes or heavier training drive frontier model gains?

Across six frontier models in RSIGym's Joint track, harness fixes consistently improved performance while additional fine-tuning lowered scores in eight of ten comparisons. Claude agents ranked highest by spending more time diagnosing and revising harnesses rather than scaling training.

Can a stronger model lift a weaker one at test time without retraining?

A stronger model built inference-time harnesses that nearly doubled weaker model performance on Theory-of-Mind benchmarks without retraining, primarily by moving unstable reasoning into deterministic code and task-specific routing rather than encouraging extended reasoning.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.