Does sharing experience across agents beat isolated evolution?
Can pooling code patches and task traces across a group of evolving agents sustain progress better than keeping lineages separate? This tests whether diversity becomes useful stepping stones or remains wasted exploration.
Group-Evolving Agents (GEA) makes "a group of agents, rather than an individual agent, as the fundamental unit of evolution." Prior open-ended self-evolving systems follow "chain- or tree-structured evolution, where different branches remain strictly isolated," so exploratory diversity "rarely serves as effective stepping stones" and instead produces "short-lived variants that fail to contribute to long-term cumulative progress." GEA's fix is explicit experience sharing within a group at each iteration. On the paper's own benchmarks this measures out to 71.0% on SWE-bench Verified and 88.3% on Polyglot, against 56.7% and 68.3% for the prior open-ended self-evolving methods it compares against — and, "without any human intervention," comparable to or beating human-designed frameworks (71.8% and 52.0% on the same two benchmarks).
Mechanically, each iteration has two stages. Parent Group Selection picks K agents from an archive using a "Performance–Novelty" score: agents are represented as task-success vectors, novelty is cosine distance to nearest neighbors, and the combined score weights performance as primary with novelty as "a mild bias." Open-Ended Group Evolution then has the parent group jointly produce a same-size offspring group by pooling each agent's evolutionary traces — code patches, a predicted patch for a sampled unsolved task, that task's execution logs and tool-invocation history, and the evaluation outcome exposing failure modes — across the whole group rather than within one lineage. The paper's own accounting of why this works: nine tool-level modifications drove the gains, GEA's best agent integrated eight of them, the comparison system's best agent integrated only five, and "the four tools missing from the DGM agent were explored in isolated branches... but failed to propagate due to lineage isolation." Five of GEA's eight originated from different parent agents — the sharing mechanism, not just more search, is doing the work.
This reads as a worked instance of the first stage in Can agents evolve beyond the constraints humans engineer? — peers exerting adaptive pressure on each other — with benchmark numbers behind the survey's claim that single-entity evolution stalls inside isolated branches. It also names its foil directly: the Darwin Gödel Machine of Can AI systems improve themselves through trial and error? is the "DGM agent" GEA out-integrates above, so GEA keeps DGM's empirical-validation-plus-archive recipe and replaces DGM's individual-lineage parent selection with group-level selection and cross-agent sharing. And because GEA's gains are, in its own analysis, "workflow and tool enhancements rather than model-specific optimizations" on frozen GPT- and Claude-series backbones, this sits inside the fast loop of Do self-improving agents really split into two distinct loops? — no weights change, only the scaffold each agent carries.
The excerpt does not name the specific prior methods behind the 56.7%/68.3% and 71.8%/52.0% comparison figures, so how directly comparable those baselines are cannot be checked from this text alone. All results come from the authors' own runs in "isolated sandbox environments," with no independent replication reported here, and the group size K, archive size, and compute budget behind these numbers aren't given in this excerpt. The authors themselves flag that open-ended exploration "may inadvertently introduce directions misaligned with human intent" and produce patches "difficult to fully understand," and their extension to non-coding domains (they mention bias mitigation) is stated as potential, not demonstrated. The defensible claim is narrower than "group evolution beats individual evolution" in general: within this paper's coding benchmarks and this archive-based evolutionary setup, cross-agent experience sharing converted diversity that would otherwise die out in isolated branches into measurable, transferable gains.
Inquiring lines that read this note 10
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can AI agents improve their skills through accumulated experience and reuse? What limits recursive self-improvement in autonomous AI systems?- How do evolutionary archives improve on single self-modification trajectories?
- How do evolutionary archives enable open-ended self-improvement without formal proofs?
- How did individual agents shift toward collective swarm behavior?
- Does co-evolution between peer agents reduce reliance on human design?
- Can pluralism survive within a single platform or does it require architectural exits?
Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can agents evolve beyond the constraints humans engineer?
Does removing human-designed elements from self-improvement systems—starting with peer agents, then task design, finally the update mechanism itself—allow artificial agents to escape the limits of static learning contexts?
GEA is a concrete peer-level co-evolution system, giving benchmark evidence for the survey's claim that single-entity self-evolution stalls
-
Can AI systems improve themselves through trial and error?
Explores whether replacing formal proof requirements with empirical benchmark testing enables AI systems to successfully modify and improve their own code iteratively, and what mechanisms prevent compounding failures.
GEA's named baseline; keeps DGM's empirical-archive recipe but replaces individual-lineage selection with group-level sharing
-
Do self-improving agents really split into two distinct loops?
Explores whether modern self-improving agents can be understood through a clean abstraction separating fast scaffold updates from slow model weight updates, and whether this framework actually explains the field's recent progress.
GEA's gains are entirely scaffold/workflow-level on frozen backbones, placing it in the fast loop
-
How can agent systems share learned skills across users?
Individual users operating autonomous agents independently rediscover solutions because systems lack mechanisms to propagate discoveries. Can centralized aggregation and automatic evolution convert isolated experiences into shared capabilities?
parallel claim that aggregating experience across otherwise-siloed agents, not individual learning alone, drives propagation of discoveries
-
Does benchmark score predict a coding agent's self-improvement capacity?
When self-improving agents are ranked by immediate coding-benchmark performance, does this reliably identify which agents will produce the most productive descendants? This matters because tree-search self-improvement relies on choosing which agent variant to expand next.
qualifies A: benchmark-point gains may not predict true self-improvement productivity, per B's metaproductivity-performance mismatch finding
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Group-Evolving Agents: Open-Ended Self-Improvement via Experience Sharing
- Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design
- DarwinX: Evolving Agent Harnesses Through Natural Selection
- Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents
- SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
- Autogenesis: A Self-Evolving Agent Protocol
- Rethinking the Evaluation of Harness Evolution for Agents
- Large Language Model Agents Are Not Always Faithful Self-Evolvers
Original note title
group-evolving agents share experience across parents instead of isolating lineages, turning transient diversity into sustained progress