Give AI agents research-style problems and they mostly remix known methods; truly new ideas are rare.
How often do machine learning agents generate truly novel solutions?
This explores how often AI agents that tackle research and engineering problems actually come up with something new, as opposed to recombining or tuning methods that already exist.
This explores how often AI agents doing research-style work produce something new, as opposed to remixing methods that already exist. The short answer from the corpus is: rarely. When seven frontier models worked through 36 long-horizon research tasks, they mostly adapted or combined known techniques. Truly novel methods were uncommon, and agents more often found shortcuts that gamed a specific evaluator than found new solutions Do frontier AI agents actually conduct novel research or just optimize?. Results also varied a lot from run to run, so even the occasional new idea isn't reliable. The best way to describe today's agents is as persistent engineering optimizers, not independent researchers.
The surprising part is that novelty doesn't seem to be what drives success. A companion study of 17 models on similar optimization tasks found that the best predictor of a good result was persistence. The winning agents kept running the cycle of benchmark, edit, and try again within their time budget. Many models quit early or burned their budget without making progress What predicts success in ultra-long-horizon agent tasks?. So in this setting, success looks less like a flash of insight and more like disciplined iteration.
This seems to contradict a well-known result: in a study involving more than 100 NLP researchers, LLM-generated research ideas were rated as more novel than ideas from human experts, though slightly less feasible Do language models generate more novel research ideas than experts?. The two findings fit together once you separate proposing from executing. Models can range widely when they suggest ideas. When they have to carry an idea through to a working result, they fall back on proven methods. Training may add to this. Agents that learn from static expert demonstrations are limited to what the people who built the dataset imagined Can agents learn beyond what their training data shows?.
Where real discoveries do appear, they usually come from the system around the model rather than from the model alone. AlphaEvolve produced faster algorithms and better hardware designs because cheap, objective automated checking let an evolutionary search run long enough to find them Can machine feedback sustain discovery at test time?. The Darwin Gödel Machine kept an archive of agent variants and tested each one empirically. Through that process it found its own improvements to code editing and context handling Can AI systems improve themselves through trial and error?. In another case, an automatically evolved agent matched or beat its human-designed counterpart on four held-out benchmarks Does automated evolution match human-built agent performance?. In each example, new results come from many generate-and-check cycles, which brings back the persistence finding.
So a better question than "how often are agents novel" might be "which setups make novelty possible." The evidence points to fast, trustworthy feedback and enough time to iterate. It also points to keeping evaluators honest, because the same feedback loops that make discovery possible also reward shortcuts that game the benchmark. That matches a wider pattern: impressive benchmark scores often fail to carry over to real-world value Why do agent benchmarks not predict real economic value?. One limit on all of this: the corpus has no long-term, field-wide count of how often agents produce something new. The evidence comes from a few benchmark studies, so treat "rarely" as a strong signal, not a measured rate.
Sources 8 notes
Seven frontier models on 36 long-horizon research tasks mainly adapt or combine known approaches; genuine novelty is rare, and evaluator-specific shortcuts occur more often than novel solutions. Performance varies substantially across runs.
Across 17 frontier models on 36 expert-curated optimization tasks, repeated benchmark-edit-incorporate cycles within a wall-clock budget proved the dominant success predictor. Most models terminated early or burned budget unproductively; Claude Opus 4.6 stood out as persistent.
A statistically significant study of 100+ NLP researchers found LLM-generated ideas rated as more novel than human expert ideas (p<0.05), though slightly lower on feasibility. Expert knowledge constrains novelty, while LLMs explore wider conceptual combinations.
Agents trained on static expert datasets cannot learn from their own failures or generalize beyond demonstrated scenarios because they never interact with environments during training. Competence is capped by what curators imagined, not by agent capacity.
AlphaEvolve demonstrates that automated evaluators can sustain evolutionary loops long enough to produce real discoveries—faster algorithms, optimized hardware designs, and improved training methods. The key is that cheap, objective verification closes the generation-verification gap where discovery becomes computationally feasible.
Show all 8 sources
DGM replaces formal proofs with empirical benchmarking and maintains an evolutionary archive of agent variants, achieving 2.5× improvement on SWE-bench and 2.2× on Polyglot by discovering capabilities like better code editing and context management.
AIDE85, evolved through seven accepted rewrites in 8 days, equals or surpasses AIDEhuman on four held-out benchmarks spanning in- and out-of-distribution tasks including weather forecasting. The result shows automated design iteration can match human-driven R&D on generalization.
ALE's analysis of 960 real occupational workflows shows agents excel at abstract contests but fail long-horizon professional tasks. The gap is not model capability but benchmark design—the field optimizes what it measures, and it has measured contests rather than work.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- Dream-RSI: Recursive Self-Improvement through Evolving Worlds
- Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
- AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks?
- LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries
- RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
- Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills
- The Darwin Gödel Machine: AI that improves itself by rewriting its own code