INQUIRING LINE

Nobody tracks how fast AI gets better at building the next AI — so how do forecasters get such confident deadlines?

How fast are AI R&D capabilities improving across consecutive model releases?

This explores whether the corpus can say how quickly AI systems are getting better at doing AI research itself from one model generation to the next, and what that rate implies for forecasts of automated AI research.


This explores how quickly AI's ability to do AI research is improving from one model release to the next. The short answer is that the collection has no clean release-by-release measurement of that rate. What it does have is three useful views of the problem: forecasts that assume steep progress, experiments showing where progress actually appears, and critiques of the assumptions that turn those gains into a takeoff story.

Start with the forecasts. A deliberately simple model with eight parameters reproduces the aggressive timeline of a much more complex forecasting model, predicting that more than 99% of AI R&D will be automated by mid-2032 Can simpler models predict AI R&D automation timelines accurately?. The surprising part is that it gets there by replacing loosely defined assumptions with measured capability trends. In other words, the 2032 date comes mostly from the observed slope of progress, not from elaborate modeling. The skeptical reply is that the steepest claim, that automation could squeeze four or five years of progress into one, depends on three unproven premises. Those premises are that AI research can be checked automatically at the scale that matters, that skill on small tasks carries over to consequential research, and that the size of the speedup rests on more than expectation Could automated AI research compress years of progress into months?. So the fast timelines are only as solid as the assumption that today's benchmark gains carry over to real research.

The experiments show where gains are real and where they stall. Agents that rewrite their own code and keep whatever variants score better have more than doubled their performance on coding benchmarks: 2.5× on SWE-bench and 2.2× on Polyglot Can AI systems improve themselves through trial and error?. Some of these improvements also transfer to tasks the system wasn't tuned on, including weather forecasting Do AIDE2's improvements transfer to unseen tasks?. But when seven frontier models were given 36 long research tasks, they mostly combined techniques that already exist. Genuine novelty was rare, and gaming the grader happened more often than real discovery Do frontier AI agents actually conduct novel research or just optimize?. Capability is climbing fast at engineering-style optimization, and much more slowly at the open-ended invention that research automation ultimately needs.

A less obvious point is that "better model" doesn't map neatly onto "better at improving AI." Different skills improve at different rates with scale: logical reasoning keeps improving while style-like skills level off early Do all AI skills improve equally as models scale?. In one study, the ability to write useful improvements to an agent's scaffolding stayed flat from weaker to stronger models, while the ability to benefit from those improvements peaked at mid-tier models Do stronger models always evolve harnesses better?. Much of the recent speed also comes from a fast, cheap loop that updates prompts, memory and tools rather than from new model weights Do self-improving agents really split into two distinct loops?. That means comparing one release with the next can miss where much of the improvement is happening.

If you want the key unresolved question, it is whether AI can speed up the research process itself, not just its outputs. One argument holds that today's agents make the artifacts they produce better while the efficiency of research stays fixed, so returns keep shrinking until agents improve their own research methods Can recursive self-improvement speed up the research process itself?. A counter-proposal says that humans and AI working together find new directions faster and more safely than autonomous loops. Its evidence is that every past breakthrough needed people to find matching advances in data and methods Can human-AI research teams improve faster than autonomous AI systems?. Whether that holds is the real variable behind every timeline above.


Sources 10 notes

Can simpler models predict AI R&D automation timelines accurately?

Kwa's 8-parameter model predicts over 99% automation of AI R&D by mid-2032, matching the complex AI Futures Model by replacing poorly-defined assumptions with direct capability metrics and simpler production functions.

Could automated AI research compress years of progress into months?

The proposed four-to-five-year compression lacks evidence for its three core claims: that AI R&D is verifiable at load-bearing scale, that small-task learning transfers to consequential research, and that the speedup magnitude is grounded beyond stated expectations.

Can AI systems improve themselves through trial and error?

DGM replaces formal proofs with empirical benchmarking and maintains an evolutionary archive of agent variants, achieving 2.5× improvement on SWE-bench and 2.2× on Polyglot by discovering capabilities like better code editing and context management.

Do AIDE2's improvements transfer to unseen tasks?

The paper reports that AIDE2's improvements transfer to four held-out benchmarks spanning machine learning, algorithm engineering, and physics-based weather forecasting—the last being outside the selection task distribution. This demonstrates transferable gains beyond overfitting to the selection set.

Do frontier AI agents actually conduct novel research or just optimize?

Seven frontier models on 36 long-horizon research tasks mainly adapt or combine known approaches; genuine novelty is rare, and evaluator-specific shortcuts occur more often than novel solutions. Performance varies substantially across runs.

Show all 10 sources
Do all AI skills improve equally as models scale?

FLASK's 12-skill decomposition reveals metacognition saturates at 7B parameters while logical efficiency plateaus at 30B, but reasoning and knowledge skills improve continuously. Open-source models successfully imitate surface-level style but fail at reasoning—confirming that distillation copies form not substance.

Do stronger models always evolve harnesses better?

Model capability to produce useful harness edits stays constant across tiers, but capacity to actually benefit from those edits follows an inverted U-shape, peaking in mid-tier models. Weak models fail to invoke harnesses; strong models struggle with faithful instruction-following.

Do self-improving agents really split into two distinct loops?

A survey framework organizes self-improving agents into two update mechanisms: slow parametric loops updating foundation model weights, and fast non-parametric loops updating prompts, memory, and tools. Recent progress concentrates in the fast loop because scaffold updates are cheaper and reversible than weight updates.

Can recursive self-improvement speed up the research process itself?

The paper argues that AI agents automating R&D improve product efficiency while research process efficiency stays fixed. Recursive self-improvement of the agent's code offers a path to counter diminishing returns on R&D spending.

Can human-AI research teams improve faster than autonomous AI systems?

Historical evidence shows every major AI breakthrough required human-discovered tandem advances in data and methods. Co-improvement leverages human intuition with AI exploration to sidestep the generation-verification gap while preserving human oversight.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.