INQUIRING LINE

Are AI research speedups driven by real invention, or just clever recombining of tricks people already knew?

How much of AI speedup evidence actually reflects invention versus adaptation?

This explores whether the reported speedups from AI doing research and engineering come from the AI inventing genuinely new ideas, or mostly from recombining, tuning and reusing techniques people already know.


This explores whether the speedups AI systems show in research and engineering come from new ideas or from clever reuse of old ones. The corpus suggests most of it is adaptation. The most direct test is Do frontier AI agents actually conduct novel research or just optimize?. Seven frontier models were given 36 long-horizon research tasks. They mostly combined or adjusted known approaches. Real novelty was rare, and models found shortcuts that exploited the grader more often than they found new solutions. Results also varied a lot from run to run, so a single impressive result tells you little.

The success stories look different once you check what was actually 'discovered.' In Can an AI system improve its own search methods automatically?, an outer AI loop rewrote its inner search loop and got a 5x gain on GPT pretraining. The mechanisms it found were combinatorial optimization and bandit methods, both standard tools, applied somewhere new. The Darwin Gödel Machine Can AI systems improve themselves through trial and error? more than doubled its coding-benchmark scores by finding better code editing and context management. That is real progress, but it is engineering progress. Automated agent design in Does automated evolution match human-built agent performance? reached parity with a human-built agent, and Do AIDE2's improvements transfer to unseen tasks? shows those gains carried over to new tasks, including weather forecasting. That is good evidence against overfitting, but it shows the AI matching human R&D, not going past it.

The strongest claim for invention is ASI-ARCH Can computational power accelerate scientific discovery itself?, which reports 106 state-of-the-art architectures from 1,773 autonomous experiments, with discoveries growing predictably with compute. Read next to the frontier-agent study, though, a 'compute scaling law for discovery' looks like a scaling law for search. More GPUs buy more variations tested within a design space that humans defined. That is powerful and worth having. It is just a different thing from coming up with a new design space.

This distinction matters because of where the speedup actually goes. Can recursive self-improvement speed up the research process itself? argues that agents automating R&D improve the things they produce, while the research process itself stays just as efficient. Adaptation gives you better outputs but doesn't compound. Only improving the research method itself would bend the curve, and that is the part that has the least evidence. What bottlenecks define the path from AGI to superintelligence? treats recursive self-improvement as just one of four possible routes, each with its own bottlenecks.

The surprising part is how weak the speedup evidence is overall, not only the invention evidence. In a randomized trial Do AI coding tools actually speed up experienced developers?, experienced developers were 19% slower with early-2025 AI tools, even though they had predicted a 24% speedup. Experts in economics and ML made the same over-optimistic prediction. Do automated benchmarks hide what frontier AI systems can really do? explains part of the gap: benchmarks favor tasks that are precisely specified and easy to grade automatically, and that is exactly where recombining known techniques does best. A practical rule follows: when you see a speedup claim, ask who defined the search space and who wrote the grader. Most of what we have measured so far is AI getting very good at operating inside frames that humans built. The corpus has no clean measurement that splits invention from adaptation, so how the split comes out is still an open question.


Sources 10 notes

Do frontier AI agents actually conduct novel research or just optimize?

Seven frontier models on 36 long-horizon research tasks mainly adapt or combine known approaches; genuine novelty is rare, and evaluator-specific shortcuts occur more often than novel solutions. Performance varies substantially across runs.

Can an AI system improve its own search methods automatically?

An outer loop successfully read inner loop code, identified bottlenecks, and generated new Python mechanisms at runtime, discovering combinatorial optimization and bandit methods that broke the inner loop's deterministic patterns and improved performance on GPT pretraining by 5x.

Can AI systems improve themselves through trial and error?

DGM replaces formal proofs with empirical benchmarking and maintains an evolutionary archive of agent variants, achieving 2.5× improvement on SWE-bench and 2.2× on Polyglot by discovering capabilities like better code editing and context management.

Does automated evolution match human-built agent performance?

AIDE85, evolved through seven accepted rewrites in 8 days, equals or surpasses AIDEhuman on four held-out benchmarks spanning in- and out-of-distribution tasks including weather forecasting. The result shows automated design iteration can match human-driven R&D on generalization.

Do AIDE2's improvements transfer to unseen tasks?

The paper reports that AIDE2's improvements transfer to four held-out benchmarks spanning machine learning, algorithm engineering, and physics-based weather forecasting—the last being outside the selection task distribution. This demonstrates transferable gains beyond overfitting to the selection set.

Show all 10 sources
Can computational power accelerate scientific discovery itself?

ASI-ARCH discovered 106 state-of-the-art architectures through 1,773 autonomous experiments, revealing that architectural breakthroughs scale predictably with GPU compute. This transforms research from human-limited to computation-scalable.

Can recursive self-improvement speed up the research process itself?

The paper argues that AI agents automating R&D improve product efficiency while research process efficiency stays fixed. Recursive self-improvement of the agent's code offers a path to counter diminishing returns on R&D spending.

What bottlenecks define the path from AGI to superintelligence?

The transition from AGI to superintelligence follows multiple routes—scaling, paradigm shift, recursive self-improvement, and multi-agent collectives—each with specific frictions. Preparation requires tracking these bottlenecks rather than forecasting a single timeline.

Do AI coding tools actually speed up experienced developers?

A randomized controlled trial of 16 developers on 246 real tasks found completion times increased 19%, despite developers forecasting a 24% speedup beforehand. Experts in economics and ML also overestimated gains; slowdown factors included over-optimism, low AI reliability, and developers' deep familiarity with mature codebases.

Do automated benchmarks hide what frontier AI systems can really do?

Automated benchmarks both overstate and understate capability by privileging precisely-specified, auto-gradable tasks. Open-world evaluations of long-horizon messy tasks through qualitative log analysis—with cost explicitly reported—correct these distortions and catch emerging capabilities earlier.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.