Does recursive self-improvement start with harness engineering?
Explores whether near-term RSI advances through optimizing deployment systems and orchestration layers rather than models directly rewriting their own weights, and what evidence supports this pathway.
Weng argues the recursive self-improvement (RSI) loop in modern AI need not take the form Yudkowsky described — "an AI uses its current intelligence to improve the cognitive machinery that produces its intelligence" — as a model rewriting its own weights. The feedback loop, she writes, "may indicate the model rewriting its own weights directly, or more broadly the model improves the training pipeline and the deployment system," and she treats the deployment layer, the harness, as "as important as the model's raw intelligence." Her prediction is explicit: "the near-term path of RSI is unlikely to start as a model directly rewriting its weights" — it instead runs through harness engineering, the system that "orchestrates execution and decides how the model thinks and plans, calls tools and acts, perceives and manages context, stores artifacts, and evaluates results."
The mechanism she proposes is a staged progression of what gets optimized: "instruction prompts → structured context → workflow → harness code → optimizer code." As models grow more capable, the target of improvement climbs this ladder toward more generic, less heuristic mechanisms — "the harness system itself becomes an optimization target, with fewer heuristic rules and more general mechanisms." She draws an explicit analogy to operating systems (a harness should "encapsulate complicated logic while keeping the interface simple") and to the history of prompt engineering: "manual prompt tricks became less central as instruction tuning and model reasoning improved, but the need to specify goals, constraints, context, and evaluation did not disappear." By that analogy she expects harness-level gains eventually to be "internalized into core model behavior," while the interface to external context and tools persists.
This reframes several library notes as evidence of rungs on the progression Weng describes, rather than as isolated results. Can language models build and maintain their own agent harnesses? supplies the empirical basis for her "deployment system" framing: HarnessDev shows the same model scoring differently under different harnesses, which is exactly the separation her argument needs. Does self-editing through reviewed commits improve agent performance? is a concrete instance of her "harness code" rung, with human review standing in for the heuristic rules she expects eventually to fall away. Can a routing harness generate its own training data automatically? makes a parallel bet that RSI's near-term engine is the deployed harness layer rather than direct weight modification. Can harness modules improve separately from benchmark data? targets the generalization problem her progression implies: a harness optimized on one benchmark is stuck on a lower rung unless its updates are built to transfer.
The excerpt is a position essay, not a measurement: Weng cites other groups' benchmark results (AHE on Terminal-Bench-2, Self-Harness on Terminal-Bench-2) in passing but reports no data of her own, and her claim that harness gains will eventually be "internalized into core model behavior" is a stated prediction, not something the post demonstrates. It also does not specify what would count as a harness update crossing into weight-level RSI, leaving the boundary between "the deployment system improves" and "the model improves itself" to the reader's judgment. The implication, held at that strength, is that progress toward RSI is for now better tracked by which rung of the stack — prompts, context, workflow, harness code, optimizer code — a given result moves, than by whether any single benchmark score went up.
Inquiring lines that read this note 17
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can AI research automation sustain progress through accelerating feedback loops? What limits recursive self-improvement in autonomous AI systems?- Why do cybersecurity and self-improvement capability thresholds move at different rates?
- How does OpenAI's Preparedness Framework define AI self-improvement capability?
- What minimum model capability is required before self-improvement bootstrapping can begin?
- How do single-improver systems compare to population-based self-improvement?
- What failure modes does recursive self-improvement encounter in evolutionary loops?
- Can scaffold-only refinement scale to open-ended recursive self-improvement?
- How does bounded self-refinement differ from open-ended recursive self-improvement?
- How do recursive self-improvement and iterative policy improvement differ fundamentally?
- What distinguishes bounded self-refinement from open-ended recursive self-improvement?
- How much token efficiency can harness-level intervention achieve versus training approaches?
- Can harnesses that rewrite themselves through reviewed commits achieve state-of-the-art performance?
- Can routing harnesses contain the mechanisms needed for deployed recursive self-improvement?
Related concepts in this collection 7
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can language models build and maintain their own agent harnesses?
This explores whether an LLM's ability to create and revise its own execution infrastructure is a distinct skill from solving tasks within someone else's harness, and whether current evaluations overlook this capability.
supplies the empirical separation Weng's "deployment system is as important as the model" claim depends on
-
Does self-editing through reviewed commits improve agent performance?
Ouroboros evolves its own prompts, tools, and core code through a reviewed commit process and reports top benchmark scores. But without comparing to a frozen version of itself, the contribution of self-evolution versus initial design or model capacity remains unclear.
a concrete instance of the "harness code" rung in Weng's predicted optimization progression
-
Can a routing harness generate its own training data automatically?
Explores whether the logs and signals produced by an agent routing system—which directs requests to appropriate model tiers—naturally contain the evidence needed to improve the models themselves through fine-tuning and distillation.
shares Weng's bet that RSI's near-term engine is the deployed harness, not weight rewriting
-
Can harness modules improve separately from benchmark data?
Does evolving harness components independently on out-of-distribution data, using contrasted success and failure trajectories, help distinguish reusable improvements from task-specific overfitting? This matters because current methods conflate general gains with benchmark adaptation.
addresses the generalization problem implied by Weng's harness-code rung
-
Do harness fixes or heavier training drive frontier model gains?
RSIGym tested whether improving model weights through additional training or refining the evaluation harness itself produces better performance across six frontier models. Understanding which lever works matters for allocating research effort and budget.
evidence for Weng: RSIGym trials show harness revisions drove gains while heavier training lowered scores in most comparisons
-
Can agent harnesses be automatically optimized across many environments?
Explores whether scaling auto-research loops across diverse harness environments can discover mechanisms that reduce token use without sacrificing task performance, and whether such discoveries generalize.
evidence for Weng: scaled auto-research loops found harness-level mechanisms cutting token traffic nearly half with no weight changes
-
Can sequencing edits across three surfaces avoid hidden costs?
When improving deployed models, editing data, harness, or weights separately each incurs invisible costs to the others. Does running them in a planned sequence with shared signals prevent those costs from accumulating?
qualifies Weng: MetaRSI-v1 needs scheduled harness-data-model sequencing, not harness alone, since each surface pays a blind cost
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement
- Harness Engineering for Self-Improvement
- Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops
- MetaRSI / RSI2: A Meta-Recursive Self-Improving System for Recursive Self-Improving Systems Themselves
- ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement
- NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness
- Dream-RSI: Recursive Self-Improvement through Evolving Worlds
- The Economics of Recursive Self-Improvement
Original note title
Weng argues RSI's near-term path runs through harness engineering, not models rewriting their own weights