INQUIRING LINE

Why do AI 'evolution' experiments miss good tricks found by other branches that never got shared?

Why do isolated evolutionary branches fail to propagate useful discoveries?

This explores why AI systems that improve themselves by evolving separate lineages of agents (each branch building only on its own parent) miss useful discoveries that other branches make, and what the corpus says about fixing this.


This explores why self-improving agent systems that evolve in separate, walled-off lineages miss out on good ideas, and what changes when lineages share. The clearest evidence comes from Group-Evolving Agents Does sharing experience across agents beat isolated evolution?. In the usual tree-style setup, each child agent inherits only from its own parent, so a clever tool fix found on one branch never reaches the others. When the researchers pooled code patches and execution traces across every agent in a generation, performance rose by 14–20 percentage points. The detail that matters most: five of the eight key tool improvements came from a *different* parent than the agent that ended up using them. Useful discoveries tend to show up on one branch and pay off on another, and isolation cuts exactly that connection.

One reason isolation hurts is that a lineage feeding only on its own output runs into the same limits as any closed self-improvement loop. The corpus's broader account of self-improvement Can models reliably improve themselves without external feedback? argues that pure self-improvement stalls through diversity collapse and reward hacking, and that methods that work bring in outside reference points. To any single branch, the other branches act as that outside reference point: they offer variation it could not have produced alone. A related failure happens inside a single model's reasoning. Reasoning models often drop promising solution paths too early and wander instead Why do reasoning models abandon promising solution paths?. Without a shared record, what one attempt learned is simply lost.

What the corpus adds that you might not expect is that sharing failures matters as much as sharing successes. AutoScientists Can decentralized teams outperform central planners in long-running science? beat centralized planners on biomedical tasks under the same experiment budget. It did this with self-organizing teams that kept competing hypotheses alive *and* passed around what didn't work. AutoResearchClaw's pivot-or-refine loop Can experiment failures drive progress instead of stopping it? makes the same point inside one agent: a failure becomes a signal for the next attempt rather than a dead end. An isolated branch that fails keeps that lesson to itself, so the other branches may repeat the mistake.

Sharing doesn't happen on its own, though, and it can get worse as systems grow. Work on scaling agent populations Does scaling agent populations thin mutual observation? finds that as populations grow, each member sees less of what the others are doing and its link to the group weakens. Adding branches can therefore make isolation worse. Some systems answer this with deliberate structure. SkillClaw How can agent systems share learned skills across users? collects trajectories from many users in one place, refines skills from them, and pushes the updates out to everyone. Dream-RSI Can past discoveries train better exploration policies? goes further and treats the whole tree of past discoveries as a cheap simulator for training better exploration strategies. In that view, a branch's history is shared data to replay later, not just a record of its ancestry.

The short answer is that isolated branches don't fail because each one searches badly. They fail because a discovery's value often shows up somewhere other than where it was found, and isolation keeps it from getting there. Classic evolutionary search can rediscover deep ideas from scratch Can evolutionary search discover machine learning algorithms from scratch?, but rediscovering things independently is expensive. A pooled memory lets the population skip work it has already done.


Sources 9 notes

Does sharing experience across agents beat isolated evolution?

Group-Evolving Agents outperformed isolated tree-based self-evolution by 14–20 percentage points by explicitly pooling code patches and execution traces within each generation. Analysis showed five of eight key tool improvements came from different parent agents, proving the sharing mechanism itself—not just more search—drove the gains.

Can models reliably improve themselves without external feedback?

Pure self-improvement stalls due to the generation-verification gap, diversity collapse, and reward hacking. Reliable improvement methods succeed by smuggling in external anchors: past model versions, third-party judges, user corrections, or tool feedback.

Why do reasoning models abandon promising solution paths?

Reasoning LLMs exhibit two reinforcing failures: wandering (invalid exploration) and underthinking (premature path-switching). Decoding-level interventions like thought-switching penalties improve accuracy without fine-tuning, suggesting viable solutions exist but are abandoned prematurely.

Can decentralized teams outperform central planners in long-running science?

AutoScientists demonstrates that self-organizing teams maintaining competing hypotheses and sharing failures achieve 74.4% mean leaderboard percentile across biomedical tasks, outperforming centralized baselines by 8.33% under matched experimental budgets.

Can experiment failures drive progress instead of stopping it?

AutoResearchClaw's pivot-or-refine loop routes every failure through a decision process, making failure inform the next attempt rather than stop execution. Component ablation shows this mechanism drives completion and is distinct from reasoning or verification.

Show all 9 sources
Does scaling agent populations thin mutual observation?

Research suggests defection in scaled populations is structural, not motivational. As populations grow, components' links to the collective weaken and their observational scope shrinks, reducing the visibility that enforces norm compliance.

How can agent systems share learned skills across users?

SkillClaw aggregates interaction trajectories across users, processes them through an autonomous evolver that identifies patterns and refines skills, then synchronizes updates system-wide. This converts siloed individual learning into shared capability improvement without manual curation.

Can past discoveries train better exploration policies?

Dream-RSI demonstrates that accumulated discovery trees can be replayed off-policy to score exploration policies without repeated online evaluation. The framework loops between policy evaluation on historical data, online redeployment, and simulator expansion, reportedly achieving competitive discovery quality at lower cost.

Can evolutionary search discover machine learning algorithms from scratch?

AutoML-Zero evolved algorithms from 65 basic operations that match neural networks and rediscover modern techniques like weight averaging and learning-rate decay, adapting strategies to task conditions in controlled experiments.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.