INQUIRING LINE

Could the system that just picks which AI model handles a task already be the seed of an AI that upgrades itself?

Can routing harnesses contain the mechanisms needed for deployed recursive self-improvement?

This explores whether the software layer that decides which model handles each request (a routing harness) already has what it takes to let a deployed AI system keep improving itself, and what would still be missing.


This explores whether a routing harness, the deployed layer that sends each request to a particular model, could serve as the engine of recursive self-improvement. The corpus gives a qualified yes: the ingredients are already there, but they add up to a loop that is real and bounded, not an open-ended takeoff. The strongest claim is that a working router is already a data factory. It logs which model got which task, how hard the task seemed, and whether the result worked. Those logs can be turned directly into labeled examples for fine-tuning and distillation. In effect the router becomes a difficulty labeler as a side effect of doing its job Can a routing harness generate its own training data automatically?. This fits Weng's argument that the near-term path to self-improvement runs through harness engineering (prompts, then harness code, then optimizer code) rather than models rewriting their own weights Does recursive self-improvement start with harness engineering?.

The less obvious reason this might work is the one that makes most self-improvement schemes fail. Pure self-improvement tends to go in circles. Models can't reliably check their own work, their outputs become less varied, and they learn to game their own rewards. The methods that do work succeed by quietly bringing in an outside anchor, such as user corrections, tool feedback, or a separate judge Can models reliably improve themselves without external feedback?. A deployed router has that anchor built in, because its outcome data comes from real tasks succeeding or failing in the world. A related idea treats an accumulated history of discoveries as a cheap replay simulator for testing new exploration strategies Can past discoveries train better exploration policies?. Router logs could play the same role, letting a system try new routing policies on past traffic before using them live.

It helps to think of self-improving agents as running two loops: a fast loop that updates prompts, memory, and tools, and a slow loop that updates model weights Do self-improving agents really split into two distinct loops?. A routing harness sits where the two meet. Its routing decisions belong to the fast loop, and its logged trajectories feed the slow one. The harness side alone has real room to improve. Automated harness tuning across many environments found four mechanisms that cut token traffic by nearly half without losing performance Can agent harnesses be automatically optimized across many environments?. Reorganizing a harness around how the code behaves at runtime let weaker planners match stronger models at finding the right code to change Can explicit behavior maps help weaker planners compete with stronger models?. Both results matter to a router, because it can then send more work to cheaper models. Making the router itself improvable is easier when its parts are standardized; one framework breaks every router into five comparable components Can five components unify all LLM routing approaches?, and another uses versioned capability profiles so routing doesn't have to be wired by hand Can semantic capability vectors replace manual agent routing?.

There are three catches. First, the loop may not compound the way people expect. Models of every size are about equally good at proposing harness edits, but the benefit from those edits peaks at mid-tier models. Weak models fail to use the harness at all, and strong models struggle to follow its instructions faithfully Do stronger models always evolve harnesses better?. So a stronger model trained on router data could get less out of the router that produced that data. Second, what a router supports is bounded self-refinement: measurable, improvement on a known task mix. A 1,250-paper survey argues that this is a different phenomenon from open-ended recursive self-improvement, which is still limited by the need for grounding, by collapse dynamics, and by compute Are self-refinement and recursive self-improvement actually the same thing?. Third, whether any of these loops accelerates depends on the product of how strongly each feedback pathway responds. By current estimates that product is growing but is still too small to keep acceleration going on its own Are AI feedback loops strong enough to sustain recursive self-improvement?. The router supplies a mechanism; whether that mechanism compounds is still an open question.


Sources 12 notes

Can a routing harness generate its own training data automatically?

A deployed routing system records execution trajectories, capability demand estimates, and outcome data that can be converted into labeled training examples for fine-tuning and distillation, turning the harness into both a serving component and a difficulty labeler.

Does recursive self-improvement start with harness engineering?

Weng argues RSI's initial path moves through optimizing deployment harnesses—instruction prompts to harness code to optimizer code—rather than models directly rewriting weights. This staged progression mirrors how prompt engineering gave way to instruction tuning while interface needs persisted.

Can models reliably improve themselves without external feedback?

Pure self-improvement stalls due to the generation-verification gap, diversity collapse, and reward hacking. Reliable improvement methods succeed by smuggling in external anchors: past model versions, third-party judges, user corrections, or tool feedback.

Can past discoveries train better exploration policies?

Dream-RSI demonstrates that accumulated discovery trees can be replayed off-policy to score exploration policies without repeated online evaluation. The framework loops between policy evaluation on historical data, online redeployment, and simulator expansion, reportedly achieving competitive discovery quality at lower cost.

Do self-improving agents really split into two distinct loops?

A survey framework organizes self-improving agents into two update mechanisms: slow parametric loops updating foundation model weights, and fast non-parametric loops updating prompts, memory, and tools. Recent progress concentrates in the fast loop because scaffold updates are cheaper and reversible than weight updates.

Show all 12 sources
Can agent harnesses be automatically optimized across many environments?

Cross-environment harness optimization yielded four mechanisms (action execution, context compaction, observation handling, delegated reading) that reduced token traffic by 44.7–49.0% while maintaining comparable performance on a 51-task benchmark, suggesting harness-level gains are orthogonal to model improvements.

Can explicit behavior maps help weaker planners compete with stronger models?

A behavior-to-code mapping representation improved win rates by 10–19 points while reducing planner tokens by 8–13%. Weaker planners using this mapping matched stronger models' code localization across all precision and recall metrics.

Can five components unify all LLM routing approaches?

The LLMRouter framework casts routing as a sequential decision process with five component types, enabling fair comparison of diverse routers and unifying single-turn, multi-turn, and personalized routing as instances of a common design space.

Can semantic capability vectors replace manual agent routing?

Versioned Capability Vectors embedded in HNSW indices couple semantic matching with policy and budget constraints, making capability discovery a first-class operation that scales sub-linearly as agent heterogeneity increases.

Do stronger models always evolve harnesses better?

Model capability to produce useful harness edits stays constant across tiers, but capacity to actually benefit from those edits follows an inverted U-shape, peaking in mid-tier models. Weak models fail to invoke harnesses; strong models struggle with faithful instruction-following.

Are self-refinement and recursive self-improvement actually the same thing?

A 1,250-paper survey shows that bounded, evaluable self-refinement (current industrial practice) differs fundamentally from open-ended recursive self-improvement, which remains constrained by grounding requirements, collapse dynamics, and compute limits measurable today.

Are AI feedback loops strong enough to sustain recursive self-improvement?

Back-of-the-envelope modeling shows recursive improvement loops depend on the product of elasticities across feedback pathways. Current loops remain too weak for self-sustaining acceleration, though they appear to be strengthening based on data on researcher productivity and system benchmarking trends.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.