If AI starts writing and improving its own code, how fast could that snowball — and when might it take over most of the work?
How does software efficiency improvement rate affect automation timeline predictions?
This explores how assumptions about the pace of software-side progress (better algorithms, harnesses and agents, as opposed to more compute) drive forecasts of when AI could automate AI research itself, and what evidence the corpus has for setting that pace.
This explores how assumptions about software progress (algorithms, scaffolding and agent design rather than more compute) shape forecasts of when AI automates its own research. The corpus doesn't have a paper that isolates the software-efficiency rate as a single variable and shows how timelines shift when it changes. What it does have is the pieces that make the rate matter: a forecasting model, evidence on how much AI actually speeds up engineers, and early demonstrations of AI improving AI software.
The clearest anchor is a stripped-down forecasting exercise. An 8-parameter model reproduces the aggressive prediction of a far more complex forecast: over 99% automation of AI R&D by around mid-2032 Can simpler models predict AI R&D automation timelines accurately?. Its contribution is methodological. It swaps poorly defined assumptions for direct capability measurements and simpler production functions. That matters because the final date depends heavily on how fast research output compounds once AI contributes to it. When the inputs are measured instead of assumed, the forecast stands or falls on how reliable those measurements are.
That is where the empirical evidence pushes back. In a randomized trial, early-2025 AI coding tools made experienced open-source developers 19% slower, even though the developers expected a 24% speedup and outside experts in economics and ML also predicted gains Do AI coding tools actually speed up experienced developers?. Any timeline model that assumes AI already speeds up software work inherits that optimism. The capability side is also uncertain: automated benchmarks can both overstate and understate what systems can do on messy, long-horizon work Do automated benchmarks hide what frontier AI systems can really do?. So the input timeline models depend on most, how fast AI improves the software that builds AI, is also one of the hardest to measure.
There is direct evidence that this kind of software progress is starting. The Darwin Gödel Machine improved its own agent code through trial and benchmarking, reaching 2.5× gains on SWE-bench Can AI systems improve themselves through trial and error?. An automatically evolved research agent matched or beat its human-built predecessor after seven accepted rewrites in 8 days Does automated evolution match human-built agent performance?, and those gains carried over to tasks outside its training set, including weather forecasting Do AIDE2's improvements transfer to unseen tasks?. Lilian Weng argues that this is where recursive self-improvement starts in the near term: AI rewriting the prompts, harness code and optimizers around a model, not its own weights Does recursive self-improvement start with harness engineering?. Efficiency gains also show up in models themselves: a compact 35B model trained on real execution trajectories sits at the low-cost end of the cost-performance frontier, so training choices can partly substitute for raw scale Does model efficiency matter more than peak capability for real work?.
The takeaway is that software progress is the part of the forecast that can compound on itself. Compute grows on hardware and investment schedules, but if AI improves the code that builds AI, the efficiency rate is no longer a fixed input. It becomes an output of the process being forecast, which is why small changes in that assumption can move a timeline by years. The corpus currently holds two contrasting data points: systems that measurably improve their own code on benchmarks, and a controlled trial where real engineers got slower. Whichever one predicts the future better will likely decide whether dates like 2032 hold.
Sources 8 notes
Kwa's 8-parameter model predicts over 99% automation of AI R&D by mid-2032, matching the complex AI Futures Model by replacing poorly-defined assumptions with direct capability metrics and simpler production functions.
A randomized controlled trial of 16 developers on 246 real tasks found completion times increased 19%, despite developers forecasting a 24% speedup beforehand. Experts in economics and ML also overestimated gains; slowdown factors included over-optimism, low AI reliability, and developers' deep familiarity with mature codebases.
Automated benchmarks both overstate and understate capability by privileging precisely-specified, auto-gradable tasks. Open-world evaluations of long-horizon messy tasks through qualitative log analysis—with cost explicitly reported—correct these distortions and catch emerging capabilities earlier.
DGM replaces formal proofs with empirical benchmarking and maintains an evolutionary archive of agent variants, achieving 2.5× improvement on SWE-bench and 2.2× on Polyglot by discovering capabilities like better code editing and context management.
AIDE85, evolved through seven accepted rewrites in 8 days, equals or surpasses AIDEhuman on four held-out benchmarks spanning in- and out-of-distribution tasks including weather forecasting. The result shows automated design iteration can match human-driven R&D on generalization.
Show all 8 sources
The paper reports that AIDE2's improvements transfer to four held-out benchmarks spanning machine learning, algorithm engineering, and physics-based weather forecasting—the last being outside the selection task distribution. This demonstrates transferable gains beyond overfitting to the selection set.
Weng argues RSI's initial path moves through optimizing deployment harnesses—instruction prompts to harness code to optimizer code—rather than models directly rewriting weights. This staged progression mirrors how prompt engineering gave way to instruction tuning while interface needs persisted.
Occamy-1.0, a 35B-parameter model further trained on execution-grounded data and long-horizon trajectories, achieves competitive performance with much larger models while sitting at the low-cost knee of the Pareto frontier, suggesting that training for coordination and follow-through substitutes for raw scale in multi-step work.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- Agents' Last Exam
- The Darwin Gödel Machine: AI that improves itself by rewriting its own code
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
- We are Changing our Developer Productivity Experiment Design
- AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks?
- RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
- Open-World Evaluations for Measuring Frontier AI Capabilities