MetaRSI / RSI2: A Meta-Recursive Self-Improving System for Recursive Self-Improving Systems Themselves

Paper · arXiv 2609.06396 · Published September 6, 2026
Frontier AI Risk & RSI

Abstract: Recursive self-improvement (RSI) lets a system improve the model-building machinery from its own failures, so every later model inherits the gain. Yet RSI has been validated almost exclusively on coding and formal benchmarks such as science QA and mathematics. This format bound limits RSI to improvement within a machine-checkable slice, not general capability where questions are open and correctness is settled by argument, replication, or measurement. We argue RSI must next operate across real, diverse scientific, engineering, and meta-scientific domains, not where formal evaluation is merely tractable. To that end we present MetaRSI-v1, where improvement is the scheduled composition of three typed operators over one unified paradigm. Data-RSI amplifies existing competence and marks its boundary; Harness-RSI edits a five-slot scaffold without touching weights; Model-RSI internalizes capability into parameters through bounded training. Sharing one loop kernel and artifact vocabulary, they make data, scaffold, and model changes composable rather than exclusive. A two-axis optimizer jointly decides operator order and each operator’s proposal policy, while a meta-level policy revises the schedule across terms. We validate MetaRSI-v1 under the field’s standard evaluations, on code and closed-form science, with no external teacher: the target model plays every role in its own loop. MetaRSI-v1 reframes self-improvement from a single-surface edit to a composition across the full model-production pipeline, opening two paths: a model route internalizing capability through training, and a harness route leaving weights untouched and thus extending self-improvement to any model reachable through an interface, with Data-RSI redefined as the shared substrate feeding both. The framework further yields refutable laws on where loops exist, how operators compose, and what supervision buys.

Introduction. 1 WHAT RSI HAS TO BECOME It has only closed where checking was free. seven in ten improve against a machine-checkable target The constraint is format, not subject. and only three in ten work on code at all So build the machinery that builds loops.

One Kernel, Two Axes, Three Operators 1. Introduction Recursive self-improvement, or RSI, asks whether a system can read its own failures and improve the model-building machinery itself, so that every model built thereafter inherits the gain. Recent systems show this is no longer speculative: agents rewrite their own scaffolds [1–3] and generate their own finetuning data [4, 5]. Yet RSI has so far been validated almost exclusively on coding tasks and formal benchmarks such as multiple-choice science QA, mathematical problem sets, and executable test suites. In our survey of forty-five recent systems, 69% close their improvement loop against a target a machine can check for free, and the systems reporting scientific benchmarks fall inside that majority rather than outside it (Figure 4). This is a format bound rather than a subject bound. What such a loop certifies is restricted, benchmark-bound capability, meaning competence inside a pre-specified, machine-checkable slice of a discipline, not general capability in the discipline itself, where the question is not given I THE EVIDENCE a deployed system runs a term what it did, and how it was judged one deficiency, and no label saying which surface THE SAME EVIDENCE the same whether the work is code, a proof, a protocol or a diagnosis II THREE SINGLE-SURFACE READINGS HARNESS ONLY reads the evidence extend the scaffold: prompt, memory, built-in tools, skills, MCP re-paid at every inference, out of a bounded budget DATA ONLY reads the evidence amplify the corpus: press out what is held, mark the boundary cannot reach capability the rollouts never exhibit MODEL ONLY reads the evidence internalize the behaviour into the weights pays a training run for what may be a scaffold rule III WHAT NONE OF THEM CAN SAY one example of a composed sequence H extend the scaffold D amplify what that revealed M internalize it into the weights H retire the extension ô capability retained, scaffold budget returned scaffold budget paid at every inference returned the answer is a sequence across all three surfaces harness surface data surface model surface the cost this operator cannot see ô retired after internalization Figure 3: Why single-surface improvement is insufficient. I: a deficiency produces evidence that carries no label saying which of the three surfaces is responsible. II: each single-surface operator closes a valid loop on its own surface but incurs a cost it cannot see (red): scaffold extensions are re-paid as context at every inference, data amplification cannot reach capability the rollouts never exhibit, and training may regress already-held behaviour. III: a composed sequence pays neither. Extend the scaffold, amplify what it reveals, internalize into weights, then retire the extension: this retains the capability while returning the scaffold budget to baseline. The chart’s ordinate is that budget, so the fall at station four is the cost returned. Colours denote the three surfaces: data, harness, and the model; the scissors mark licensed retirement. and correctness is settled by argument, replication, or measurement. Nor can specialization defer the problem: a deployed system’s behaviour is determined jointly by its data, its weights, and its execution scaffold, and a change in any one alters what the other two describe (Figure 3), so what is needed is a unified paradigm over all three surfaces together with a meta-level policy governing which change is made next. We present MetaRSI-v1, a meta-recursive self-improving system built on three typed operators that share one loop kernel and consume the same learning signal. Data-RSI synthesizes verified training records from execution experience, amplifying what the model already does well and marking where that competence ends. Harness-RSI edits the execution scaffold through typed patches over five slots: the system prompt, memory, built-in tools, skills, and tools and resources mounted through MCP, taking effect at once with no training cost. Model-RSI modifies model parameters and architecture under bounded training recipes, internalizing capability into the weights so that it persists across scaffolds and adds no inference cost. The operators are connected by Transition Agent-v1 adapters, which pass the learning signal and their outputs between them. Above them, the RSI2 Agent-v1 optimizes on two axes. Horizontally, it decides the sequence in which operators are applied, since the order changes what each subsequent operator reads and some orders are ill-posed; vertically, it rewrites each operator’s own proposal policy, improving how that operator diagnoses failures and proposes changes.

Related work. 2.1. Foundation Models The capability that self-improvement now attempts to extend was itself produced by a sequence of engineering regimes, and the shape of that sequence explains where the remaining headroom lies. Scale was the first regime: the Transformer [6] made compute the binding constraint, and empirical scaling laws turned model size, data volume, and compute into a predictable trade [7, 8], yielding models whose few-shot competence emerged without task-specific training [9]. Alignment was the second: instruction tuning and reinforcement learning from human feedback converted raw next-token competence into a followable interface [10, 11], and open-weight families made that interface broadly reproducible [12–14]. Inference-time computation was the third: chain-of-thought prompting showed that additional serial computation buys accuracy at fixed weights [15], a trade later made explicit by test-time scaling analyses [16] and internalized by reasoning models trained with verifiable rewards [17, 18]. The fourth regime, and the one that current systems occupy, moved capability out of the weights entirely: retrieval, tools, memory, and long-horizon agent loops [19–21] now mediate most of what a deployed model can accomplish, and standardized interfaces for tool and context provision have made this scaffold a first-class engineering artifact. Read as a sequence, these regimes show a steady migration of the effective locus of capability: from parameters, to alignment data, to inference procedure, to the surrounding execution structure. A self-improvement framework restricted to any one of them therefore addresses only part of what determines a model’s behavior. That same migration is why MetaRSI-v1 treats data, scaffold, and the model as three writable surfaces of one system rather than as three separate targets.

2.2. Recursive Self-Improvement The idea is old, driving the classical intelligence-explosion argument [22] and formalized in the Gödel machine [23], but only recently buildable. What has been built sorts by which part of the system it treats as mutable, and that sorting is what makes the field’s concentration visible. The earliest work treats the output: self-refinement and verbal reflection revise an answer within an episode [24, 25], which is cheap to check but leaves nothing behind and, without an external signal, corrects little [26]. A second family treats the scaffold: STOP improves the program that improves programs [1]; ADAS and AFlow search over architectures and workflow graphs [2, 27]; Gödel Agent and the Darwin Gödel Machine rewrite their own code under empirical selection [3, 28]; and Self-Harness, DemoEvolve and MetaSkill-Evolve target harness evolution directly [29–31].

Method. 3. Preliminaries This section fixes the objects that any self-improvement procedure manipulates, the signal it is allowed to read, and the criterion against which its output is judged. All notation introduced here is reused unchanged in the remainder of the paper.

Target system The entity being improved is not a checkpoint but a triple S = (D, θ, H), (1) where D is the data state (the corpus, synthesized records, and curriculum over which the model has been or will be trained), θ is the model state (trainable parameters and adapters, the bounded architectural choices that place them, and the training configuration that produced them), and H is the harness state, the execution scaffold that mediates every interaction with the model, comprising the system prompt, a persistent memory, a built-in tool set, a skill library, and the tools and resources mounted through MCP. A deployed system is the composition of all three; a change confined to any one of them changes what the deployed system does.

Tasks, rollouts, and verification Let T denote a task family and t ∼T a task instance. Executing S on t produces a trajectory τ = (o1, a1, . . . , on, an) of observations and actions, terminating in an output that a verifier v maps to a scalar outcome r = v(t, τ) ∈[0, 1]. We deliberately do not assume that v is cheap: in executable domains it is a test suite, in closed-form question answering a string comparison, and in open-ended scientific work it is expensive, partial, or unavailable. This variation, rather than any property of the underlying reasoning, is what has historically determined where improvement loops could be closed.

Learning signal A self-improvement procedure never observes T directly. It observes a learning signal σ = Σ(E), E = (tk, τk, yk) K k=1, (2) the deterministic compilation of an evidence set E drawn from attempts made by the system in its current state, where yk is the outcome record of attempt k: in the code and QA tracks evaluated here, simply the verifier scalar rk above, and in general whatever the domain was able to establish about that attempt. We define σ by the role it plays rather than by the form it takes, because across domains the form varies and the role does not. In implementation σ is the compiled evidence bundle carrying raw trajectories and diagnoses; later references to trajectories or experience denote components of this bundle.

What may enter the evidence set Four kinds of thing: verifier outcomes of any fidelity, the system’s own trajectories (read for the competence they exhibit, not only for whether they succeeded), external knowledge admitted against a demonstrated gap, and the framework’s own decision records. A source is admissible if it derives from the current system, is readable by every operator, and carries no authority to decide whether a change was good. What makes something a learning signal is its position in the loop, not its modality.

What the compiler emits Σ is deterministic and fixed: it aggregates, attributes, and attaches provenance, emitting one typed object in three layers (an outcome layer, an attribution layer, and a model-attributed mechanism layer, whose schema is Equation (10)), and it does not decide what to change. Fixing the typing while leaving the sources open is what lets a single operator set run over domains whose evidence has nothing else in common.

Two properties of the signal It is the sole channel between the environment and any improvement procedure, which is what allows heterogeneous procedures to be compared, exchanged, and composed. And it is perishable: a step that changes what the deployed system does invalidates every signal compiled before it, because the attempts such a signal summarizes were made by a system that no longer exists.

Discussion. Section 1 argued that the object self-improvement should act on is the human labour that turns compute into a deployed model, and that current systems automate the cheapest slice of it. This section returns to that claim and asks what the framework actually buys against it: first for one deployment (Section 6.1), then for the closed feedback loop the three operators form and the two routes through it (Section 6.3), then across scientific domains (Section 6.4), and finally at the boundary where a loop stops being digital and starts touching apparatus (Section 6.5).

6.1. What Composition Buys Three things change when improvement is expressed as a scheduled composition of typed operators rather than as a loop around one editable surface, and each corresponds to one of the three mechanisms identified in Section 1.

Against misattribution A single-surface operator’s diagnosis is constrained by its write access; under the loop kernel it is constrained by the evidence instead, since the failure signature of Equation (10) is grounded before any operator sees it and its vocabulary is read by all three. So a missing procedural habit routes to the harness and absent knowledge routes to data and weights, on a shared object no operator can rewrite in its own favour. The learning signature of Equation (11) sharpens this on the data side: only the capability the rollouts never exhibit justifies spending external supervision.

Against non-composability The engineer’s answer from Section 1 (add the rule, collect the corrected behaviour, internalize it, delete the rule) is the path H →D →M →H of Figure 8, every arrow a typed adapter. What the framework adds is that the path is sayable: a scheduler can propose it, a type-check can accept it, and its final step can be replayed and reverted if the deletion turns out to cost accuracy.

Against measurement in the cheapest regime The framework requires a verifier of some fidelity, plus the three separations of Section 4.7, which are what let a loop be trusted when the verifier is weak: they keep the loop from improving its own definition of success. A domain with a partial verifier is therefore a harder instance of the same problem rather than a different one, and Data-RSI supplies what such a domain additionally needs, a bounded statement of where external supervision is actually required (Section 6.4).

6.2. An Amplifier Is Not a Source The composition of the three operators has a fixed upper bound. Data-RSI renders explicit the competence latent in the model’s rollouts; Model-RSI fixes that competence in parameters; Harness-RSI makes it available at inference without training. Each redistributes ability the model already holds. A record synthesized for a capability the model has never exhibited is authored by the same model that lacks it (Section 4.2.1), so the loop cannot bootstrap knowledge it does not contain. The system composed of these operators alone is a closed amplifier, whose ceiling is the best arrangement of what was already present. Every genuine gain requires information originating outside the loop. The framework admits such information through the learning signal, which is defined by its position in the loop and imposes no restriction on its source (Section 3). A verifier’s ruling, a retrieved document, a curated corpus, an instrument reading, and a commissioned expert note enter on equal terms. Human supervision occupies the same footing: the labour that self-improvement is said to displace is, in this accounting, one information source distinguished primarily by its cost. Data-RSI governs the expenditure of that cost.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

What limits recursive self-improvement in autonomous AI systems? What are the fundamental limits of prompting for language models? Why does AI verification capability persistently exceed generation capability?