When do AI agents outperform human research experts?
RE-Bench tests whether AI agents can match or exceed human ML researchers on open-ended engineering tasks across different time budgets, to understand at what scale each excels.
RE-Bench (Research Engineering Benchmark, V1) is METR's direct comparison of AI agents against human ML-research experts, drawing on "71 8-hour attempts by 61 distinct human experts" across "7 challenging, open-ended ML research engineering environments." The paper reports a time-budget crossover: "the best AI agents achieve a score 4× higher than human experts when both are given a total time budget of 2 hours," but "humans currently display better returns to increasing time budgets, narrowly exceeding the top AI agent scores given an 8-hour budget, and achieving 2× the score of the top AI agent when both are given 32 total hours (across different attempts)." The human baseline itself is nontrivial: "82% of expert attempts achieving a non-zero score and 24% matching or exceeding our strong reference solutions."
The design is built to make that comparison fair and resistant to early saturation: each environment supplies a scoring function the agent can query at any time, a weak starting solution, and a hidden high-scoring reference solution used only for normalization, with the whole suite constrained to runs human experts can meaningfully progress on in 8 hours with at most 6 H100s. Agents are scored via best-of-k across varying time budgets rather than a single run, which is how the paper surfaces the crossover rather than a flat win/loss number. Qualitatively, "most agent solutions score close to 0," but agents "submit new solutions over 10 times faster than human experts, and occasionally find very successful approaches" — both o1-preview and Claude 3.5 Sonnet found kernel-optimization solutions that "beat the efforts of all 9 human experts." Cost favors agents sharply: "~$123" per 8-hour agent run against "approximately $1,855" paid per human expert.
This gives a crossover and cost/speed reading that Do frontier AI agents actually conduct novel research or just optimize? doesn't have: that paper's within-run metrics explain why a final score can hide research capability, but it carries no human baseline and no time-budget axis, where RE-Bench supplies exactly that axis (at the cost of covering only 7 hand-built environments against that paper's 36 tasks and seven models). The Triton kernel win is also a partial exception to that paper's finding that strong agent solutions "mainly adapt or combine established techniques" — RE-Bench's own qualitative read treats the kernel result as a genuinely novel approach, though it flags this as occasional, not typical. Separately, "fixed evaluation budget" means something different here than in Do fixed-budget efficiency gains translate to real research progress?: there it holds compute fixed inside a self-improvement loop; here the budget is the wall-clock time varied to compare humans against agents.
The excerpt is explicit that this is a narrow proxy: the environments are "cleanly defined, non-interacting tasks" with no coordination between workstreams, "at least 2 orders of magnitude" smaller in scale and complexity than real frontier AI R&D, and built without the scaffold iteration the authors expect would raise agent scores further. The paper's own discussion argues, rather than measures, that "the human–AI gap in real-world AI R&D" is likely "much larger than observed on these evaluations," and that agents matching top human RE-Bench performance "may still be far from capable of AI R&D automation." At the strength the data allows: as of this benchmark (November 2024), AI agents already win decisively on short, cheap, fast-iteration R&D work, and whether closing the longer-horizon gap is mainly a scaffolding problem (which the authors expect to narrow it) or a problem of task novelty and coordination (which this data cannot speak to) is left open.
Inquiring lines that read this note 17
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Why do autonomous agents misreport success on failed actions? How do real-world evaluations reveal AI capabilities that benchmarks hide?- How do lab-scale benchmark tasks differ from real frontier AI research?
- What gap exists between AI model capability in benchmarks and real client work?
- Can frontier AI models match expert human performance on specialized tasks?
- How much faster and cheaper are AI agents compared to human researchers?
- Do AI agents and human researchers follow the same optimization patterns?
- Do research agents mostly reproduce known techniques or discover novel solutions?
- Can we distinguish agent effort from actual research output quality?
- Can agentic AI systems handle judgment-intensive tasks in science?
Related concepts in this collection 2
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do frontier AI agents actually conduct novel research or just optimize?
Exploring whether current long-horizon research agents generate genuine methodological novelty or primarily recombine established techniques. This matters for understanding how close we are to recursive self-improvement through AI.
same score-hides-capability concern, but this paper adds the human-baseline time-budget crossover that one lacks
-
Do fixed-budget efficiency gains translate to real research progress?
The paper measures research efficiency as optimization gains under a fixed evaluation budget, but this differs from the real-world costs of R&D spending and human effort. Does this narrower measurement actually predict whether AI agents reduce the true cost of research discovery?
a different sense of "fixed budget": there it fixes compute inside a self-improvement loop, here it varies wall-clock time to compare humans against agents
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
- Expenditure Horizon: Measuring Optimization Ability, with an Application to NanoGPT
- MatrAIx: Simulating the World with 8.3 Billion Persona Agents
- Recursive self-improvement of AI research agents
- The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
- Atria Dawn: The Dawn of Agentic Superintelligence
Original note title
METR's RE-Bench finds AI agents score 4x human experts at a 2-hour budget but humans narrowly exceed agents at 8 hours and lead 2x at 32