SYNTHESIS NOTE
Topics›Novel Architectures›this note

Can AI systems improve themselves through trial and error?

Explores whether replacing formal proof requirements with empirical benchmark testing enables AI systems to successfully modify and improve their own code iteratively, and what mechanisms prevent compounding failures.

Synthesis note · 2026-02-23 · sourced from Novel Architectures

The original Gödel Machine proposed self-improving AI via provably beneficial self-modifications. In practice, formally proving the impact of most self-modifications is impossible. The Darwin Gödel Machine (DGM) replaces formal proofs with empirical validation: try modifications, test them on benchmarks, keep what works. This mirrors biological evolution — mutations are not verified in advance but produced, trialed, and selected.

DGM alternates between self-modification and evaluation phases. During self-modification, agents from the archive generate modified versions of themselves — rewriting their own code. During evaluation, each modified agent is tested on coding benchmarks. The key assumption: improvement on coding benchmarks indicates better coding capabilities, which in turn indicates better ability to self-modify. This creates a meta-competence loop: better coding → better self-modification → better coding.

Results: SWE-bench from 20.0% to 50.0%, Polyglot from 14.2% to 30.7%.

The evolutionary archive is critical. Inspired by open-endedness research, DGM maintains a growing library of all generated agent variants — including suboptimal but interesting ones. These serve as stepping stones for future generations, enabling diverse exploration paths. The system doesn't just optimize for immediate performance; it accumulates diverse capabilities that may enable future breakthroughs. This is fundamentally different from single-trajectory self-improvement.

Concrete improvements discovered include better code editing tools, long-context window management, and peer-review mechanisms — capabilities the original agent lacked that emerged through the self-improvement process.

The Python-based implementation makes the self-modification space Turing-complete in principle. The current version modifies agent design (tools, workflows) with frozen foundation models. Full self-improvement — rewriting training scripts, training new foundation models — is left as future work.

This directly addresses What limits how much models can improve themselves?: DGM circumvents the formal proof requirement by using empirical validation, but inherits a different limitation — improvement is bounded by what the benchmark can measure. The archive approach partially addresses How quickly do errors compound during model self-training? by maintaining diverse populations rather than following single improvement trajectories.

Inquiring lines that read this note 218

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can AI systems achieve real improvement without external human feedback? Why does polished AI output gain credibility despite fundamental verifiability problems? Why do confident AI outputs mislead human trust calibration? Should models ask for clarification when facing ambiguous or under-specified information? Why do standard evaluation practices obscure safety-critical AI failures? How should systems validate code that agents generate? What limits recursive self-improvement in autonomous AI systems? Can AI systems discover fundamental improvements to their own architectures? How do models learn from self-generated outputs without cascading failures? Why does AI verification capability persistently exceed generation capability? When do multi-agent systems improve over single frontier models? Can inference-time computation adaptively substitute for static model capacity? Can external verification systems adequately replace learned reasoning in AI outputs? Why does self-revision amplify confidence in wrong model answers? How does diversity prevent model convergence on superficial patterns? Do individually safe AI actions create unsafe outcomes in integrated systems? When does parallel reasoning outperform sequential reasoning with the same token budget? How does fine-tuning trade off accuracy against reasoning quality? How do educators verify student capability when AI can produce indistinguishable work? Do single-axis benchmarks accurately measure agent capability for real deployment? How should agents coordinate through shared persistent code artifacts? How do real-world evaluations reveal AI capabilities that benchmarks hide? Do AI coding tools measurably improve developer productivity and code quality? What human oversight must AI research systems have? What prevents LLMs from applying their reasoning knowledge to improve outputs? Why do autonomous agents misreport success on failed actions? How can evaluations be made robust against model reward hacking? Does augmenting symbolic reasoning improve LLM logical reasoning ability? Does preference optimization undermine conversational grounding in language models? How does awareness of evaluation context influence model behavior? Can code harness improvements rival direct model scaling for capability? How do training data quality and composition affect downstream model performance? Can AI agents improve their skills through accumulated experience and reuse? How much of agent capability comes from harness versus the model itself? Can minimal training unlock latent reasoning already present in base models? How do AI systems determine and balance multiple competing objectives? How do reward signal properties affect model reasoning and safety? Do evolved harnesses learn transferable strategies or task-specific optimization artifacts? What explains the gap between benchmark scores and true reasoning capability? Should governance of agentic AI systems be runtime or design-time? Does AI-assisted research sacrifice exploration breadth for productivity gains? Can AI research automation sustain progress through accelerating feedback loops? Can AI systems evade safety evaluations through reasoning manipulation? How do users confuse explanation quality with actual system accuracy? Should GUI agents use structured screen representations instead of end-to-end vision? How do agents learn to distinguish valuable feedback from noise? How do evaluation environment design choices affect AI security? What governance mechanisms can effectively constrain widely deployed AI systems? Can we trust AI-generated mathematical proofs without understanding them? Are AI-generated articles systematically disadvantaged in search ranking and user engagement?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
27 direct connections · 184 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

darwin godel machine achieves open-ended self-improvement by replacing formal proofs with empirical validation and evolutionary archives