Can AI systems improve themselves through trial and error?
Explores whether replacing formal proof requirements with empirical benchmark testing enables AI systems to successfully modify and improve their own code iteratively, and what mechanisms prevent compounding failures.
The original Gödel Machine proposed self-improving AI via provably beneficial self-modifications. In practice, formally proving the impact of most self-modifications is impossible. The Darwin Gödel Machine (DGM) replaces formal proofs with empirical validation: try modifications, test them on benchmarks, keep what works. This mirrors biological evolution — mutations are not verified in advance but produced, trialed, and selected.
DGM alternates between self-modification and evaluation phases. During self-modification, agents from the archive generate modified versions of themselves — rewriting their own code. During evaluation, each modified agent is tested on coding benchmarks. The key assumption: improvement on coding benchmarks indicates better coding capabilities, which in turn indicates better ability to self-modify. This creates a meta-competence loop: better coding → better self-modification → better coding.
Results: SWE-bench from 20.0% to 50.0%, Polyglot from 14.2% to 30.7%.
The evolutionary archive is critical. Inspired by open-endedness research, DGM maintains a growing library of all generated agent variants — including suboptimal but interesting ones. These serve as stepping stones for future generations, enabling diverse exploration paths. The system doesn't just optimize for immediate performance; it accumulates diverse capabilities that may enable future breakthroughs. This is fundamentally different from single-trajectory self-improvement.
Concrete improvements discovered include better code editing tools, long-context window management, and peer-review mechanisms — capabilities the original agent lacked that emerged through the self-improvement process.
The Python-based implementation makes the self-modification space Turing-complete in principle. The current version modifies agent design (tools, workflows) with frozen foundation models. Full self-improvement — rewriting training scripts, training new foundation models — is left as future work.
This directly addresses What limits how much models can improve themselves?: DGM circumvents the formal proof requirement by using empirical validation, but inherits a different limitation — improvement is bounded by what the benchmark can measure. The archive approach partially addresses How quickly do errors compound during model self-training? by maintaining diverse populations rather than following single improvement trajectories.
Inquiring lines that read this note 218
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can AI systems achieve real improvement without external human feedback?- What separates performative behavioral change from actual capability development in AI?
- Can AI output be genuinely novel or only at the margins?
- What makes self-modifying architectures learn their own update rules?
- Does the 78-demonstration principle apply to other AI capabilities beyond agency?
- Can AI systems improve themselves without external feedback?
- Can AI systems design and improve their own successors without human direction?
- Can AI systems be fully understood before deployment at scale?
- Can AI models learn tacit procedural knowledge that exists only in laboratory practice?
- Can AI output be verified without understanding the reasoning behind it?
- Does verification of AI outputs face the same circularity problem?
- How does validation skill replace production skill in AI systems?
- Can AI systems produce genuinely new validity claims without community participation?
- Can expert validation scale fast enough to back AI token production?
- What makes a hypothesis match count as validation of an AI system?
- What design principles prevent error cascades in multi-step evaluation systems?
- Can AI outputs inspire new directions even when they seem like failures?
- What does recovery look like as a formal part of AI design?
- Why haven't labs adopted self-hacking approaches to catch specification errors?
- Can orchestration platforms and better infrastructure reduce AI correction time?
- Can a proposer agent actively surface a solver's weaknesses to prevent plateau?
- How should harness infrastructure validate code that agents generate themselves?
- Can skill validation through testing prevent unreliable programs from accumulating?
- How do skills authored in-loop validate faster than offline generated skills?
- Does AIDE2 archive rejected variants the way evolutionary approaches do for future reuse?
- Why do plausible edits fail when applied to running executable systems?
- Why do checkpoints get evaluated more often than actual improvements are retained?
- How do diagnose-and-reshape loops compare to building new environments from scratch?
- Can AI agents self-correct using multimodal tools to improve deliverables?
- How do agents revise their own errors during autonomous architecture discovery?
- How do evolutionary archives enable diverse exploration in self-improving systems?
- Can population diversity in self-improvement prevent error avalanching failures?
- Why do monolithic systems resist autonomous optimization attempts?
- Can co-evolved critics truly circumvent static evaluator limitations in self-improvement?
- Can multiple verification approaches together overcome the self-improvement ceiling?
- Can a model evaluate its own improvements without degrading over iterations?
- How does diversity collapse during iterative self-improvement affect solution quality?
- How does domain shift expose failures in fixed self-improvement mechanisms?
- What distinguishes iterative query refinement from pure self-revision loops?
- What four domain properties make self-healing failure loops actually work?
- Does human-in-the-loop AI collaboration accelerate recursive self-improvement safely?
- Why do most self-improving systems fail when given tasks with no clear external benchmark?
- Does removing static external utility break the formal guarantees of self-improvement loops?
- How do epoch boundaries preserve self-improvement guarantees across objective changes?
- How many acceptable rewrites can recursive self-improvement sustain before returns diminish?
- How do hidden evaluations and out-of-distribution benchmarks address recursive self-improvement risks?
- What separates a compounding improvement loop from a one-way data pipeline?
- Do evolutionary discovery systems like FunSearch count as bounded or open-ended improvement?
- Does swapping formal proofs for benchmarks change self-improvement safety?
- Why does the generation-verification gap limit what an agent can improve about itself?
- Do evolutionary archives let agents improve themselves without formal proof?
- How fast is recursive self-improvement advancing in current AI systems?
- Do diminishing returns prevent recursive self-improvement in AI systems?
- When do diminishing returns appear in repeated cycles of AI self-optimization?
- What distinguishes bounded self-refinement from open-ended recursive self-improvement in AI systems?
- What distinguishes bounded self-refinement from open-ended recursive self-improvement empirically?
- Can AI systems improve themselves through recursive self-improvement loops?
- Do bounded self-refinement and open-ended recursion pose different risk profiles?
- Does weak exogenous anchoring like compilation checks suffice for safe self-improvement?
- How does OpenAI's Preparedness Framework define AI self-improvement capability?
- Why does moving constraint descriptions change recursive improvement outcomes?
- Does weakening a verifier reduce self-improvement frequency measurably?
- What failure modes does recursive self-improvement encounter in evolutionary loops?
- Can scaffold-only refinement scale to open-ended recursive self-improvement?
- How do evolutionary archives enable open-ended self-improvement without formal proofs?
- Can archive-and-select loops sustain improvement beyond single iterations?
- How does bounded self-refinement differ from open-ended recursive self-improvement?
- Can empirical validation replace formal proofs in self-improving systems?
- What distinguishes bounded self-refinement from open-ended recursive self-improvement?
- Why do major AI breakthroughs require human-discovered data and method combinations?
- What test-time strategies did o3 discover without human specification?
- Can bilevel autoresearch autonomously modify its own learning algorithms?
- Why do good algorithms become rarer as the search space grows more generic?
- Can AI systems learn their own objectives through autoresearch?
- Why do error avalanches accelerate in self-training loops without verification?
- Can self-consistency checks fully prevent error avalanching in self-training loops?
- What causes code quality to degrade across multiple rounds of recursive self-training?
- Why do method-level improvements avoid the generation-verification gap that parameter-level improvements face?
- How does the generation-verification gap limit AI self-improvement capabilities?
- Can AI evaluation tools solve the verification problem they help create?
- How does the expert demonstration ceiling compare to the generation-verification gap bound?
- Why does human validation become the bottleneck when AI generation scales?
- Does internalizing verifiers actually close the generation-verification gap?
- Does the generation-verification gap actually limit self-improvement in verifiable tasks?
- Why does AI generation outpace verification across the research lifecycle?
- Can automated tools close the gap between AI generation and verification?
- Does the generation-verification gap limit how far AI can improve itself?
- Where does the generation-verification gap appear in test-time compute?
- How does the generation-verification gap limit autonomous discovery?
- What structural changes help AI generation keep pace with verification?
- Why does verification of AI work consistently lag behind AI generation?
- How do cheap evaluators like verifiers change discovery versus optimization?
- How does the generation-verification gap shape evolutionary program search?
- How does the verifier gap limit AI capability across different knowledge domains?
- What is the generation-verification gap that bounds self-improvement?
- What distinguishes collective evolution from vertical self-improvement in agent systems?
- Can autonomous agents detect increasingly sophisticated specification gaming as they improve?
- Can external verifiers replace reasoning trace quality in solution guarantees?
- What infrastructure could replace search for verifying AI outputs?
- Which code verification tasks still require execution instead of reasoning?
- What makes code inspectable feedback more reliable than natural language verification?
- What separates verifiable reasoning from open-ended judgment in scaling requirements?
- Can formal verifiers convert statistical semantic claims into deterministic guarantees?
- Can machines verify that formal statements mean what we intend?
- Why does most refinement in iterative models maintain answers rather than improve them?
- How does symbolic solver feedback differ from language-based self-critique?
- Can self-critique combined with integrity checks bound the self-refutation loop?
- Why do evolutionary algorithms collapse to single solutions under selection pressure?
- Can accelerated sampling techniques from image generation speed up evolutionary search?
- Can evolutionary approaches avoid the overthinking failure mode of iterative refinement?
- Why does iterative refinement fail when information stays constant?
- Can evolutionary search unlock problems that best-of-n selection cannot solve?
- Which AI safety problems lack the scalar metrics autoresearch requires?
- How should safeguards be built into AI research pipelines?
- Does population-based evolution transcend the parallel versus sequential compute tradeoff?
- Does evolutionary inference transcend the parallel versus sequential test-time compute tradeoff?
- How does low verifiability change what we can measure in AI work?
- How should we audit AI systems when transparency tools don't work as promised?
- Can systems that revise their own evaluation criteria be reliably verified?
- Can automated evaluation replace human judgment in agent testing?
- How do single-axis benchmarks misrepresent AI agent readiness for deployment?
- How do fixed external benchmarks anchor self-improving agent systems?
- Does the metaproductivity mismatch occur outside coding-agent benchmarking tasks?
- Why do a-priori procedural specifications fail as environments change and interfaces evolve?
- How should human oversight apply to persistent agent-authored code?
- Are durable shared code artifacts better than per-task harness patches?
- Why do benchmark scores not capture the true nature of AI systems?
- Can open-world evaluations become a scalable paradigm without becoming the next benchmark trap?
- Can automated benchmarks accurately capture progress on real-world long-horizon tasks?
- How do existing evaluations measure AI capability in contained environments?
- Can optimization metrics hide actual versus apparent progress in AI systems?
- Can self-administered surveys establish trustworthy AI capability benchmarks?
- How do lab-scale benchmark tasks differ from real frontier AI research?
- How do narrow benchmark optimizations differ from genuine architectural research discoveries?
- Can expenditure-matched benchmarks prevent status-driven gaming of AI metrics?
- Why does AI code generation lag behind pattern-matching benchmarks?
- What makes intermediate primitives matter more than final code execution success?
- Why do bug fixes carry more weight than hyperparameter tuning in pipelines?
- Does high-level design work benefit differently from AI than routine coding tasks?
- Does AI help more on small greenfield projects than mature codebases?
- Why does evaluation of novel primitives require waiting for future reuse to show value?
- How does validating code differ from writing code as a learning mechanism?
- Why haven't AI agents replaced human code review workflows?
- What specific failure modes appear when AI tackles research-level experiments?
- Should AI research tools separate model judgment from deterministic experiment checks?
- When should domain experts verify AI research claims before publication?
- What deterministic checks prevent AI research systems from publishing unsound claims?
- What specific research-debugging tasks measure AI self-improvement capability?
- What distinguishes verifiable AI research domains from open-ended scientific questions?
- Why do AI agents fail at verification but succeed at generation?
- Why do AI agents struggle with novel experiments but excel at routine tasks?
- What makes evaluation tamper-proof enough for autonomous research systems?
- Can trustworthy scoring prevent persistent iteration from compounding errors?
- Can an automated evaluator stay useful while an optimizer runs thousands of iterations?
- Can test environments reliably predict how models behave in actual deployment?
- What methodological shifts does model-centric evaluation require from artifact-centric testing?
- What makes API-based scaffolding more trustworthy than direct model access in high-stakes domains?
- How does AI system design amplify model capabilities beyond the weights?
- How much of AI improvement comes from tools versus model capability?
- Should we train the evolver or the executor when building self-improving agents?
- How should AI skills be created and managed like software artifacts?
- Can skill repositories evolve toward execution-oriented refinement over time?
- How does externalizing reasoning into harness artifacts improve agent reliability?
- What makes durable code artifacts more valuable than per-task harness patches?
- How does compiling natural language goals into executable code enable objective evolution?
- Can AI systems generate and refine their own objective functions?
- What persistent failures remain unsolved despite harness evolution efforts?
- How much of harness-evolution gain comes from matched test-time search budgets?
- How does test-time search budget compare to evolution gains under matched conditions?
- Can harness evolution gains be distinguished from test-time search improvements on matched budgets?
- Can harnesses that rewrite themselves through reviewed commits achieve state-of-the-art performance?
- Can empirical validation sustain long-term optimization without becoming gamed?
- Which specific AI R&D tasks does AIDE2 benchmark itself against during selection?
- How often do planted shortcuts fool autonomous research systems?
- How often do machine learning agents generate truly novel solutions?
- Is idea quality or execution capacity the actual bottleneck in AI research?
- Does human-AI collaboration improve faster and safer than autonomous self-improvement?
- How fast do new benchmarks get adopted across the AI research community?
- How should superintelligent AI systems be aligned during rapid capability gains?
- Do efficiency gains in AI-assisted development stem from better tools or autonomous improvement?
- What makes compute allocation a verifiable lever for pacing frontier AI development?
- Can recursive feedback loops turn AI research automation into genuine progress?
- What domains allow autonomous AI discovery because verification is fast enough?
- How does software efficiency improvement rate affect automation timeline predictions?
- How fast are AI R&D capabilities improving across consecutive model releases?
- How much of AI speedup evidence actually reflects invention versus adaptation?
- Can partial automation in software research alone trigger runaway AI progress?
- How does capability divergence from reliability affect AI deployment timelines?
- Can self-assessed design quality validate the actual value of AI-assisted designs?
- Can independent validation of AI output substitute for method disclosure?
- Can evaluation environments themselves become attack surfaces for AI systems?
- Does the AI Act's pre-deployment testing duty extend to post-deployment output?
- What distinguishes empirical scoring from formal proof in discovery validation?
- Can formal verification certify a proof without human comprehension?
- Does formal verification preserve human mathematical understanding across automation?
- Why do AI systems excel at literature search but struggle with novel proofs?
- Do proof assistants and neural networks fail in complementary ways?
- Can validated approximate solutions become exact mathematical proofs?
- How does this AI proof approach differ from empirical validation used in machine learning?
- Can formal proof systems eliminate the gap between checking and auditing?
- Does verification by inspection scale for AI mathematics discoveries?
- What verification methods can prove AI mathematical proofs are sound?
- Does AI-assisted research hollow out the understanding that producing proofs generates?
Related concepts in this collection 6
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
What limits how much models can improve themselves?
Explores whether self-improvement has fundamental boundaries set by how well models can verify versus generate solutions, and what this means across different task types.
DGM replaces formal verification with empirical validation, trading theoretical guarantees for practical progress
-
How quickly do errors compound during model self-training?
When LLMs train on their own outputs without verification, do small mistakes amplify exponentially? This matters because it determines whether unsupervised self-improvement is even feasible.
DGM's evolutionary archive avoids single-trajectory failure by maintaining population diversity
-
Can language models improve themselves without any external training data?
Explores whether two language models playing against each other—one generating questions, one solving them—can create a self-improving loop. Matters because it would eliminate dependence on human-labeled datasets.
related: both create self-improvement loops, but DGM modifies code rather than generating training data
-
Does learning to reward hack cause emergent misalignment in agents?
When RL agents learn reward hacking strategies in production environments, do they spontaneously develop misaligned behaviors like alignment faking and code sabotage? Understanding this could reveal how narrow deceptive behaviors generalize to broader misalignment.
DGM's benchmark-based validation is vulnerable to the same Goodhart's Law: optimizing benchmark performance may not generalize
-
Can reinforcement learning scale beyond single-turn language tasks?
Most RL for LLMs targets simple single-turn problems. This research asks whether RL can handle multi-turn interactive environments with sparse rewards and rich environmental feedback, like real software engineering tasks.
complementary path: SWE-RL achieves 39% SWE-bench via RL training on a frozen model, DGM achieves 50% via evolutionary code self-modification; suggests the combination — RL-trained agents undergoing evolutionary self-modification — could be more powerful than either alone
-
Can machine feedback sustain discovery at test time?
Can LLMs paired with automated evaluators discover genuinely novel solutions through iterative refinement, rather than just generating hypotheses? This matters because it tests whether autonomous research scales beyond benchmarks to real deployed innovations.
extends: AlphaEvolve applies the same evolutionary-archive + empirical-validation recipe to produce *deployed* algorithms (data-center scheduling, faster matrix multiply) rather than self-modifications
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- The Darwin Gödel Machine: AI that improves itself by rewriting its own code
- Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents
- Hyperagents
- Huxley-Gödel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine
- Introducing Sakana AI's Recursive Self-Improvement (RSI) Lab
- The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators
- Self-Authored Verification Is Unreliable in Heuristic Self-Improving Agents
- Self-Improvements in Modern Agentic Systems: A Survey
Original note title
darwin godel machine achieves open-ended self-improvement by replacing formal proofs with empirical validation and evolutionary archives