Can a learning strategy that evolution discovers on one problem actually carry over and still work on a totally different one?
Can evolved algorithms transfer learning strategies across different datasets and tasks?
This explores whether algorithms or agent strategies found by evolutionary search (programs mutated and selected over many generations) keep working on datasets and tasks they weren't evolved on, or whether they only fit the problem they were bred for.
This explores whether things discovered by evolutionary search, such as learning rules, programs or agent setups, carry over to new data and tasks or stay tied to the problem they were evolved on. The short answer from the corpus is yes, sometimes. But it splits transfer into two kinds. One is whether the evolving *process* travels. The other is whether the evolved *product* does. Most of the strongest evidence is about the process.
The most direct case for the product is Can evolutionary search discover machine learning algorithms from scratch?. Starting from only 65 basic math operations, evolution rediscovered neural networks, gradient descent and techniques like learning-rate decay and weight averaging. When the researchers changed the task conditions, the evolved algorithms changed their strategies to match. That shows evolution can find general-purpose learning ideas rather than one-off tricks. It also hints at a limit: what evolves reflects the conditions it evolved under. A related check comes from aidE2s-gains-generalize-to-four-held-out-benchmarks-including-physics-based-weat. AIDE2's improvements held up on four benchmarks it was never tuned on, including physics-based weather forecasting, which sits well outside its selection tasks. That is the kind of held-out test that separates real transfer from overfitting.
The process looks even more portable. Can evolutionary search beat sampling and revision at inference time? runs evolution directly in natural language, with the LLM performing crossover and mutation. Because it never needs a task-specific formal setup, the same search method moves across planning problems unchanged. Can training and search gains add together in program evolution? goes a step further. It trains a model on the evolution moves themselves (draft, improve, debug, combine), and the learned moves and the search produce gains that add up rather than overlap. So what transfers may be *skill at evolving* more than any single evolved answer.
A less obvious finding is that transfer between *agents* matters as much as transfer between tasks. In Does sharing experience across agents beat isolated evolution?, agents that pooled code patches and execution traces across lineages beat isolated evolution by 14–20 points, and most key improvements came from a different parent than the one that used them. Harness research shows that transfer depends on who receives it. A stronger model can build a harness that nearly doubles a weaker model's score (Can a stronger model lift a weaker one at test time without retraining?). But how much a model benefits from such edits peaks at mid-tier models: weak ones fail to use them and strong ones don't follow them faithfully (Do stronger models always evolve harnesses better?).
The gap: the corpus has little on taking one evolved learning algorithm and systematically testing it across many unrelated datasets. The evidence is held-out benchmark checks and condition changes, not broad transfer studies. The survey in Can agents evolve beyond the constraints humans engineer? suggests why. Self-improvement in a fixed setting tends to stall, so lasting generality may need the tasks and environments to evolve along with the algorithm.
Sources 8 notes
AutoML-Zero evolved algorithms from 65 basic operations that match neural networks and rediscover modern techniques like weight averaging and learning-rate decay, adapting strategies to task conditions in controlled experiments.
The paper reports that AIDE2's improvements transfer to four held-out benchmarks spanning machine learning, algorithm engineering, and physics-based weather forecasting—the last being outside the selection task distribution. This demonstrates transferable gains beyond overfitting to the selection set.
Mind Evolution, an evolutionary search strategy using LLM-generated crossover and mutation with island model diversity, solves 98%+ of planning tasks and significantly outperforms best-of-N and sequential revision strategies while working directly in natural language without task formalization.
Frontis-MA1 trained a single model on four program-evolution operators (Draft, Improve, Debug, Crossover), then reused those operators in long-horizon evolutionary search. The result was complementary gains: learning and search improved performance together rather than substituting for each other, lifting Medal Average from 39.39% to 71.21% on MLE-Bench Lite.
Group-Evolving Agents outperformed isolated tree-based self-evolution by 14–20 percentage points by explicitly pooling code patches and execution traces within each generation. Analysis showed five of eight key tool improvements came from different parent agents, proving the sharing mechanism itself—not just more search—drove the gains.
Show all 8 sources
A stronger model built inference-time harnesses that nearly doubled weaker model performance on Theory-of-Mind benchmarks without retraining, primarily by moving unstable reasoning into deterministic code and task-specific routing rather than encouraging extended reasoning.
Model capability to produce useful harness edits stays constant across tiers, but capacity to actually benefit from those edits follows an inverted U-shape, peaking in mid-tier models. Weak models fail to invoke harnesses; strong models struggle with faithful instruction-following.
A survey framework organizes co-evolving systems into three stages that progressively remove human engineering: dynamic peers first, then adaptive environments and feedback, finally the evolution mechanism itself. Single-entity self-improvement stalls in static contexts; co-evolution supplies adaptive pressure across multiple components.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Evolving Deeper LLM Thinking
- Learning to Discover at Test Time
- Group-Evolving Agents: Open-Ended Self-Improvement via Experience Sharing
- Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design
- Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
- Rethinking the Evaluation of Harness Evolution for Agents
- AutoML-Zero: Evolving Machine Learning Algorithms From Scratch
- Sharpening Tax in Post-Training