INQUIRING LINE

In AI research, is the hard part dreaming up good ideas, or actually executing them well?

Is idea quality or execution capacity the actual bottleneck in AI research?

This explores whether AI-driven research is held back more by coming up with good ideas or by carrying them out (coding, running experiments, writing up), and what the corpus says about where the real constraint sits.


This explores whether AI-driven research is held back more by coming up with good ideas or by carrying them out. The corpus offers a third answer: execution is fast becoming the easy part, idea selection is starting to look automatable, and the real constraint is moving to judgment, meaning deciding what's worth optimizing and checking whether the results are actually true.

Start with execution, since that's where the evidence is strongest. One system completed the whole cycle on its own: it generated ideas, wrote code, ran experiments, drafted a paper, reviewed its own work, and produced a manuscript that passed first-round review at a machine learning workshop Can one AI system complete a full research cycle end-to-end?. Agent systems also keep improving their own engineering. The Darwin Gödel Machine rewrites its own code and keeps whatever scores better on benchmarks, which more than doubled its software-engineering performance Can AI systems improve themselves through trial and error?. AIDE2's improvements carried over to tasks it was never tuned on, including weather forecasting Do AIDE2's improvements transfer to unseen tasks?. There's a catch, though. These agents make their outputs better, but the research process that produces those outputs stays just as slow. Speeding up the process itself takes agents that redesign how they work Can recursive self-improvement speed up the research process itself?.

Idea quality looks less like a fixed human advantage than you might expect. A fine-tuned GPT-4.1 with access to past papers predicted which of two research ideas would perform better 77% of the time, while expert NLP researchers did about as well as a coin flip Can machines learn to predict which research ideas will work?. Ready-made models with no extra training also scored at chance. So judging ideas is a learnable skill, but only for a model trained on what actually happened to past ideas. Generating ideas has a similar condition: groups of AI agents with different perspectives beat a single agent only when the members have real domain expertise. Without it, the variety turns into noise Does cognitive diversity alone improve multi-agent ideation quality?.

The more surprising claim is that both 'ideas' and 'execution' miss the actual choke point. One line of work argues that for big scientific problems, the hard part is defining what to optimize, not searching faster for solutions. Any fixed goal is an incomplete stand-in for the real one, and a capable optimizer will exploit the gap Why is objective design the real bottleneck in AI discovery?. A second line argues that AI produces plausible research faster than anyone can confirm it's correct. In one analysis, 39% of failures in agent-run research came from made-up content, not from misunderstanding, and the gap is widest where novelty matters most Can AI verify research outputs as fast as it generates them?. Experimental design and self-correction are also on the list of capabilities that current benchmarks don't measure, and self-correction is the weakest of them What capabilities do AI systems need for autonomous science?.

What you might not have expected to learn: when AI drafts the paper, the paper stops being evidence that someone thought it through. A polished manuscript used to signal that real reasoning happened behind it. AI breaks that link, producing the form of intellectual work without the thinking that used to come with it Does AI separate intellectual form from the thinking behind it?. So the bottleneck in AI research may not be ideas or execution. It may be trust: knowing which goals are worth pursuing and which finished-looking results are real.


Sources 10 notes

Can one AI system complete a full research cycle end-to-end?

The AI Scientist performed ideation, coding, experiments, writing, and self-review autonomously, producing a manuscript that passed the first round at a machine learning workshop with 70% acceptance rate. Five ensemble reviewers and an area-chair model judged the output against NeurIPS guidelines.

Can AI systems improve themselves through trial and error?

DGM replaces formal proofs with empirical benchmarking and maintains an evolutionary archive of agent variants, achieving 2.5× improvement on SWE-bench and 2.2× on Polyglot by discovering capabilities like better code editing and context management.

Do AIDE2's improvements transfer to unseen tasks?

The paper reports that AIDE2's improvements transfer to four held-out benchmarks spanning machine learning, algorithm engineering, and physics-based weather forecasting—the last being outside the selection task distribution. This demonstrates transferable gains beyond overfitting to the selection set.

Can recursive self-improvement speed up the research process itself?

The paper argues that AI agents automating R&D improve product efficiency while research process efficiency stays fixed. Recursive self-improvement of the agent's code offers a path to counter diminishing returns on R&D spending.

Can machines learn to predict which research ideas will work?

A fine-tuned GPT-4.1 combined with paper retrieval reached 77% accuracy predicting which of two AI ideas performs better, beating 25 expert NLP researchers 64.4% to 48.9% on a 45-pair subset. Off-the-shelf models performed at chance level, suggesting the capability requires both retrieval and fine-tuning on historical outcomes.

Show all 10 sources
Does cognitive diversity alone improve multi-agent ideation quality?

Multi-agent teams substantially outperform solo ideation, but only when members possess genuine senior knowledge. Diverse teams without expertise underperform even a single competent agent, because cognitive stimulation without expertise triggers process losses instead of insight.

Why is objective design the real bottleneck in AI discovery?

For grand challenges, fixed objective functions are incomplete and vulnerable to reward hacking. The bottleneck is automating objective design itself—the creativity of defining what to optimize—not navigating the solution space faster.

Can AI verify research outputs as fast as it generates them?

AI can produce plausible research outputs faster than it can prove them correct or meaningful, shifting the bottleneck from authorship to verification. Evidence shows 39% of agentic research failures stem from content fabrication and 32% from retrieval failures, not comprehension—and the gap widens precisely where novelty and scientific judgment matter most.

What capabilities do AI systems need for autonomous science?

The Virtuous Machines framework identifies hypothesis generation, experimental design, data analysis, and iterative self-correction as essential for autonomous scientific research, none of which standard LLM benchmarks reliably evaluate. Self-correction poses the deepest challenge due to documented degradation in reasoning accuracy.

Does AI separate intellectual form from the thinking behind it?

Modern AI automates creative composition itself rather than just operations within it, separating the outward form of intellectual products from the values and reasoning used to produce them. This mechanism allows exchange value to float free from use value.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.