Can an AI that keeps a library of its old attempts, not just its best one, keep getting smarter round after round?
Can archive-and-select loops sustain improvement beyond single iterations?
This explores whether self-improving AI loops that keep an archive of past attempts and pick the best ones to build on can keep getting better over many rounds, rather than leveling off after the first gain.
This explores whether loops that keep a growing archive of past attempts and pick from it what to build on next can keep improving over many rounds, instead of stalling after one good step. The clearest yes in the corpus is the Darwin Gödel Machine Can AI systems improve themselves through trial and error?. It drops the old idea of proving that a change is an improvement. Instead it tests agent variants on coding benchmarks and keeps an evolutionary archive of them. The archive is what keeps the loop going: weaker variants stay in the pool as stepping stones, so the search doesn't lock onto one line of descent. The gains build on each other, roughly 2.5× on SWE-bench, and they come from capabilities the system discovered for itself, like better code editing and better context management.
Surprisingly, what an archive needs is diversity more than quality. Set reinforcement learning Can diverse mediocre traces outperform redundant expert traces? shows this outside self-improvement: a set of varied, mediocre attempts beats a set of near-identical strong ones, because whatever does the selecting needs different options to choose between. Applied to archive loops, a selector that keeps only the top scorers is likely to shrink its own search space and stall. A selector that keeps useful variety leaves room for the next jump.
A second way to keep a loop going is to improve the selector, not just what it selects. Bilevel autoresearch Can an AI system improve its own search methods automatically? adds an outer loop that reads the inner loop's code, sees where its search has fallen into fixed patterns, and writes new search mechanisms (bandit methods, combinatorial optimization) at runtime. That gave a 5× improvement. SkillOS Can a separate trained curator improve skill libraries better than frozen agents? does the same for skill libraries. A separately trained curator decides what goes into the library, and over time the library shifts from generic, wordy entries toward reusable strategies that work across tasks. Both point to the same lesson: a fixed selector eventually runs out of good moves, and a selector that learns can keep finding new ones.
The risk on the other side is that an archive accumulates its own mistakes. Bidirectional RAG Can RAG systems safely learn from their own generated answers? is a knowledge-base version of the same loop. It writes its own answers back into its retrieval corpus, but only after they pass three checks: entailment, source attribution and novelty. Without those checks, errors would be retrieved later as if they were facts. Finer-grained filtering helps here too. Judging attempts step by step catches failures that one overall score hides Does step-level confidence outperform global averaging for trace filtering?.
To answer directly: yes, these loops can keep improving past one iteration, if three things hold. The archive keeps diversity rather than only winners, the selection mechanism itself can improve, and there are checks that stop bad entries from compounding. The corpus does not yet have long-run evidence on where these loops eventually plateau. The results here cover tens of iterations, not open-ended runs.
Sources 6 notes
DGM replaces formal proofs with empirical benchmarking and maintains an evolutionary archive of agent variants, achieving 2.5× improvement on SWE-bench and 2.2× on Polyglot by discovering capabilities like better code editing and context management.
SPIRAL shifts RL reward from individual traces to sampled sets, optimizing for complementarity rather than per-trace accuracy. Diverse mediocre traces outperform redundant strong ones because aggregators need raw material to arbitrate, not confirmation.
An outer loop successfully read inner loop code, identified bottlenecks, and generated new Python mechanisms at runtime, discovering combinatorial optimization and bandit methods that broke the inner loop's deterministic patterns and improved performance on GPT pretraining by 5x.
SkillOS shows that separating a trainable curator from a frozen executor, grouped by task streams, causes skill repositories to shift from generic verbose additions toward actionable execution logic and cross-task meta-strategies. The trained curator generalizes across different executor backbones and domains.
Systems can add generated answers to their retrieval corpus when outputs pass entailment verification, source attribution checks, and novelty detection. This prevents hallucinations from polluting future retrievals while allowing genuine knowledge accumulation.
Show all 6 sources
Local step-level confidence catches reasoning breakdowns that global averaging masks and enables early stopping before traces complete. This approach achieves comparable accuracy gains to naive majority voting with far fewer generated traces, proving trace quality matters more than quantity.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents
- The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators
- Local Coherence or Global Validity? Investigating RLVR Traces in Math Domains
- Bilevel Autoresearch: Meta-Autoresearching Itself
- The Darwin Gödel Machine: AI that improves itself by rewriting its own code
- SkillOS: Learning Skill Curation for Self-Evolving Agents
- Hyperagents
- Huxley-Gödel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine