RSIGym: A Flexible Environment for Recursive Self-Improvement

Paper · arXiv 2610.10310 · Published October 7, 2026
Frontier AI Risk & RSI

ABSTRACT Recursive self-improvement requires carrying accepted changes into later improvement cycles, while studying agentproposed changes also requires substantial research infrastructure. Existing settings often leave agents to rebuild routine infrastructure or restrict exploration to individual components. We introduce RSIGym, an agent-native research environment based on Everything as a Service (EaaS). RSIGym exposes training, inference, rollout, evaluation, and sandbox execution through reusable services, with shared budget and permission controls supporting Data, Harness, and Joint improvement tracks. This design enables agents to investigate individual interventions and jointly optimize data, training settings, and execution harnesses within the same environment. We define RSI-Index as the mean fraction of the remaining performance gap closed across five benchmarks covering software engineering, terminal interaction, mathematics, scientific reasoning, and skill-based tasks. Comparing six frontier research models in independent Joint runs, Opus 5 achieves the highest RSI-Index of 0.4809 under a $500 platform-service budget per benchmark run. Its selected systems improve all five benchmarks, raising SWE-bench Verified from 17.67% to 50.33% and AIME from 31.67% to 97.78%. Additional experiments examine DSH-harness refinement, budget variation, and restricted network access, while recorded trajectories reveal how agents diagnose failures and select candidates. We open-source the full RSIGym codebase and results to support reproducibility and further research.

Introduction. Following a recent survey, we use recursive self-improvement (RSI) for a process in which accepted changes are carried into later improvement cycles; recursive progress further requires that inherited changes improve how subsequent improvements are produced [Li et al., 2026a]. Automated scientific-research systems such as The AI Scientist generate ideas, run experiments, and analyze results [Lu et al., 2024, Chen et al., 2025]. In the improvement loop studied here, an agent identifies a behavioral bottleneck, proposes an intervention, and tests whether the target system improves. Yet existing environments often leave routine training and inference infrastructure for the research agent to assemble. For example, PostTrainBench gives agents an H100 GPU and internet access but no starter code, training data, or hyperparameter configuration, so agents must build the training pipeline themselves [Rank et al., 2026]. Reusable services can remove this repeated setup work and leave agents more of their time and compute budget for investigating the target system. Reducing this engineering burden must go hand in hand with preserving the agent’s freedom to explore how different components contribute to improvement. Existing benchmarks establish useful, controlled settings by focusing on particular parts of the workflow (Table 1): optimizing bounded post-training runs and training algorithms [Rank et al., 2026, Chi et al., 2026], revising training-data strategies under a shared training stack [Meng et al., 2026], or searching harnesses around fixed actor weights [Sleiman et al., 2026]. These settings help study individual interventions, while a broader RSI environment must also support their interactions. Training data and hyperparameters determine the learned model, and the harness determines how that model reasons, uses tools, and acts in a task. Studying their co-evolution therefore calls for reusable infrastructure that agents can invoke directly, together with a configurable improvement space. The same environment should support controlled studies of individual components and joint optimization of data, training settings, and harnesses. We bring these two requirements together in RSIGym, an agent-native research environment built on Everything as a Service (EaaS). Research agents work in lightweight, CPU-only containers and conduct experiments through five reusable services. A Train Server executes LoRA fine-tuning through Tinker with optional custom losses, and a Model Server exposes base models and trained checkpoints through an OpenAI-compatible gateway. A Rollout Server provides frontier models for data generation and judging, while a Benchmark Server evaluates model–harness pairs through Harbor. A separate Sandbox Service provides E2B cloud sandboxes for data generation, testing, and other computation. Together, these services support a continuous research loop: the agent proposes a change, trains or revises a candidate, evaluates it, and uses the returned evidence to decide what to try next. Familiar interfaces, service skill documents, and asynchronous execution make this loop directly accessible to research agents. The scope of this research loop is controlled through permissions and budgets, allowing the same services to support different experimental questions. A shared authorization service checks every call against the run’s permissions and remaining budget and records its cost. In the Data track, the agent changes training data while the training recipe and harness remain fixed. In the Harness track, it revises the execution code with model weights held fixed. The Joint track opens data, training settings, and harness code together. These settings make it possible to study individual interventions and their co-evolution within a common infrastructure. Our experiments measure target-system gains within individual budgeted research runs. Whether the resulting systems improve subsequent research cycles remains to be evaluated. To measure what agents achieve within these improvement spaces, we introduce RSI-Index across SWE-bench Verified, Terminal-Bench 2.0, AIME, GPQA Diamond, and SkillsBench. The index averages the fraction of the remaining performance gap closed on each benchmark, giving all five domains equal weight. In the Joint track, six frontier research models start from the same base model and minimal harness. With a $500 platform-service budget per benchmark run, Opus 5 achieves the highest RSI-Index of 0.4809. Its selected systems improve all five benchmarks, raising SWE-bench Verified from 17.67% to 50.33% and AIME from 31.67% to 97.78%. The environment’s configurable controls also let us examine how outcomes vary with the starting harness, available budget, and access to external data.

Related work. Recursive Self-Improvement (RSI). Automated research systems such as The AI Scientist generate and test scientific ideas [Lu et al., 2024, Chen et al., 2025]. A recent survey distinguishes RSI by requiring accepted changes to be retained as the starting state of later improvement cycles; it reserves recursive progress for the stronger case in which inherited changes also improve how subsequent changes are produced [Li et al., 2026a]. An earlier formal proposal for self-improvement is Schmidhuber’s Gödel machine, which searches for provably beneficial changes to its own program [Schmidhuber, 2007]. Early methods such as AlphaGo Zero self-play and STaR show how systems can improve through self-generated experience [Silver et al., 2017, Zelikman et al., 2022]. The AI Scientist extends automated research to tasks such as coding experiments and analyzing results [Lu et al., 2024]. Recent benchmarks evaluate specialized slices of this workflow, such as containerized workspaces in MLGym [Nathani et al., 2025], autonomous fine-tuning in PostTrainBench [Rank et al., 2026], reinforcement learning in Agent2 RL-Bench [Chen et al., 2026], and training-algorithm design in AI4AI-Bench [Chi et al., 2026].

Post-Training & Harnessing Optimization. DeepSeek-R1 reports that large-scale reinforcement learning can elicit reasoning behavior and improve performance on verifiable reasoning tasks [Guo et al., 2025]. Agent performance also depends on the tools and runtime that mediate model–environment interaction. AutoHarnessBench studies harness search around fixed model weights, while RSIBench-Data studies data-centered post-training under a shared training and evaluation stack [Sleiman et al., 2026, Meng et al., 2026]. These settings isolate complementary research choices: RSIBench-Data holds the post-training stack fixed while agents revise training data; AutoHarnessBench holds actor weights fixed while agents revise the harness. As formalized in Table 1, existing environments restrict autonomous exploration to isolated niches: they either optimize model weights while fixing scaffolding, search harnesses around frozen weights, or rely on fragile hostlevel scripts. Agent performance depends on both model behavior and the interaction layer. Self-Harness studies trace-driven, model-specific harness changes; HiAgent and SWE-agent examine working-memory management and agent–computer interfaces, respectively [Zhang et al., 2026, Hu et al., 2025, Yang et al., 2024]. These studies motivate evaluating model and harness changes in the same environment, while leaving the effects of their interaction to measurement.

Method. 3. System Design In this section, we formalize the budgeted target-system improvement task supported by RSIGym and detail its architecture, organized around the Everything as a Service (EaaS) paradigm, unified budget accounting, and automated scientific auditing.

3.1. Problem Definition Research task. RSIGym distinguishes the research agent π, which designs and runs experiments, from the target system (M, h), where Mis the model and his its execution harness. Each experiment specifies an initial system (M0, h0), an evaluation protocol [E], permitted changes [C], and a platform-service budget [B]total. The agent’s task is to improve the target system’s evaluation score within these permissions and the allocated budget.

Permitted changes and research process. Depending on [C], the agent may revise training data D, training settings λ, and harness code h. Through iterative experiments under the allocated budget, it selects a final model– harness pair:

Here, πdenotes the complete research process, including candidate generation, evaluation, and selection. In harness-only experiments, M⋆= M0.

Submission, evaluation, and audit. The final pair (M⋆, h⋆) is represented by a checkpoint identifier and a harness archive, with fixed components supplied by the task. The verifier evaluates this pair under [E]and returns the mean task reward. Its score and improvement over the initial system are Separate post-run audits assess submission validity through checkpoint checks, training-data overlap analysis, and behavioral review of the research trajectory and submitted artifacts.

3.2. Everything as a Service Carrying out the process in Section 3.1 requires repeated training runs, model serving, and benchmark evaluation, each of which depends on substantial infrastructure. In RSIGym, this infrastructure sits behind services rather than inside the research agent’s workspace. The agent works in a lightweight, CPU-only container that holds its harness repository, its data, and its notes. Training, inference, evaluation, and sandbox execution all run remotely, and the agent reaches them through network interfaces (Figure 1). It does not install GPU drivers, launch distributed training, deploy models, or build evaluation environments. The platform holds the credentials for all underlying providers, while the agent holds a single platform key issued for its run. Compared with giving each agent a dedicated GPU workspace, this design consumes compute only while a service executes submitted work, rather than leaving it idle while the agent reasons. Work is not bounded by a single machine; an evaluation job, for example, runs many trials concurrently in separate sandboxes. Every research agent also works with the same infrastructure, so differences between agents reflect their research decisions rather than how well they configure an environment.

Services. Table 2 lists the interface of each service, and we describe the role of each below. The training service carries out training proposals generated during the research process in Equation (1). The agent submits training data and a training configuration, optionally with a custom loss function, and obtains a checkpoint. The current implementation performs LoRA fine-tuning [Hu et al., 2022] on Tinker [Thinking Machines Lab, 2025]. The inference service makes checkpoints callable. It serves base models and trained checkpoints through an OpenAI-compatible chat completion interface with tool calls, so a checkpoint can be driven by a harness or probed directly by the agent as soon as training finishes, without an explicit deployment step. The rollout service gives access to frontier models through the same interface, routed through LiteLLM [BerriAI, 2023]. Their outputs can serve as training data, and the models can also act as judges. The evaluation service scores a model–harness pair, as in Equation (2). It runs benchmark jobs in the configuration format of Harbor [Harbor Framework Team, 2026], with each trial in its own cloud sandbox. A job may carry the agent’s own harness as a code archive, which the service installs and runs in every trial against the model endpoint named in the job.

Discussion. 5. Analysis Section 4 reports the final scores. Here we examine two common patterns—harness improvements and the effects of training—and how research behavior differs across Claude, DeepSeek, and GPT agents. We then study a Data run 5.1. Trajectory Analysis of the Joint Runs Harness improvements address recurring execution failures. Across agents, submitted harnesses address similar failures within each benchmark. SWE harnesses check edits before patch submission; Terminal-Bench harnesses preserve shell state, recover from truncated outputs, and detect unproductive loops. AIME and GPQA harnesses aggregate multiple sampled answers, while SkillsBench harnesses discover available skills and verify outputs. These changes help the target model complete and check its work. With base weights unchanged, Opus 5’s harness changes raise the Terminal-Bench development score from 6/30 to 10/30, and DeepSeek’s raise SkillsBench from 0.0232 to 0.1137 on the same 37 tasks. Thus, part of the Joint improvement can arise from changing how existing model capabilities are used.

Heavier training does not consistently improve performance. Figure 4 compares candidates under the same harness and task set. The more heavily trained candidate scores lower in eight of ten comparisons; the other two gains are one task each. For example, Opus 5’s expanded-data update reduces Terminal-Bench from 11/30 to 9/30, and Opus 5.5’s expanded-data update reduces GPQA from 86/100 to 83/100. Several agents respond by reducing the update strength: Opus 5 on Terminal-Bench and Astra on AIME submit light updates, while Opus 5.5 deliberately keeps its GPQA and Terminal-Bench checkpoints close to base weights. These small development comparisons show why more training is not a reliable default. Training can nevertheless help: in DeepSeek’s SWE run, the trained model improves the development score from 0.344 to 0.600 under the same harness, as detailed in the third point.

Research behavior differs across model families. Claude agents spend more time diagnosing failures and revising the harness. Their five-run research sessions total 23.70–37.49 hours, and they achieve the two highest RSI-Index values. In Opus 5’s Terminal-Bench run, successive tests expose shell-lifecycle, truncation, and looping problems, which lead to targeted harness fixes. When stronger fine-tuning fails to help, it reduces the update instead of continuing to increase training. This combination of repeated diagnosis and restrained weight updates accompanies the best Terminal-Bench and AIME results. Opus 5.5 similarly favors harness changes and leads GPQA with a checkpoint close to base weights. DeepSeek combines inexpensive data generation with diagnosis of training failures. It builds the largest training sets (median 472 records per run, 4,702 on AIME) and ranks third overall. On SWE, an early checkpoint produces malformed tool calls; the later model, trained on 227 trajectories for three epochs, follows the tool interface more reliably. Under the same harness on 30 tasks with three attempts each, its development score rises from 0.344 for base weights to 0.600. The submitted system scores 0.5333 officially, the highest SWE result. This case shows a productive training path alongside the regressions above. GPT agents train earlier and end their searches sooner. Their five-run sessions total 8.69–11.85 hours, with training starting after a median of 11 minutes versus 29 for the other agents. All runs end voluntarily, so the shorter searches are not imposed by a timeout. Astra rapidly builds a simple shell loop and synthetic training pipeline on SWE, reaching 0.3567 in 1.67 research hours, but remains below Claude and DeepSeek.

Conclusion. We introduce RSIGym, an agent-native environment that combines reusable research services with shared budget and permission controls. It supports individual interventions and joint optimization of data, training settings, and harnesses. RSI-Index measures improvement across domains: under a $500 budget per benchmark run, all six research agents achieve positive overall gains, led by Opus 5 at 0.4809. Research trajectories highlight execution fixes and the choice of training update strength. Revisions to the existing DSH harness improve terminal-task performance, while larger budgets do not consistently yield gains. The offline run passes the audits without beating the initial system, and Qwen3.8’s autonomous research produces no overall improvement. These results show that the environment supports both successful improvements and the study of unsuccessful research attempts.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Does AI-assisted research sacrifice exploration breadth for productivity gains? What limits recursive self-improvement in autonomous AI systems? Can AI research automation sustain progress through accelerating feedback loops?