AI research agents now log more work hours than human researchers — so what's actually left for the humans to do?
What research tasks do agents handle versus human researchers?
This explores how research work is actually being split between AI agents and human scientists: which parts of the research process agents take on, and which parts still depend on people.
This explores how research work is being split between AI agents and human researchers in practice. The short answer is that agents have taken over the middle of the research process, meaning implementation, running experiments, revising methods, analysis and writing. Humans still hold both ends: deciding what is worth asking, and deciding what counts as done. The pattern repeats across very different settings. At OpenAI, research agents now log more workdays than human staff, but that work is concentrated in implementation, and high-level planning still needs frequent human help Are AI agents now doing more research work than humans?. A study of 769 tasks from one real research project found the same split. Agents proposed methods and carried out revisions, while humans made most final calls and steered where to explore next. Participants also said about a third of those tasks wouldn't have been feasible without the agent How should AI agents and humans divide research tasks?.
The split is partly about time, not only task type. On METR's benchmark, agents score four times higher than human experts when both get two hours. By eight hours the humans have pulled slightly ahead, and at 32 hours they lead by about 2× When do AI agents outperform human research experts?. Agents are strong at short, contained work. Humans keep improving with more time in ways agents currently don't, which helps explain why long-range planning stays with people.
The less obvious finding concerns originality. Agents can technically run the whole loop: one system went from idea to code, experiments, a written paper and self-review, and passed the first round of review at a machine learning workshop Can one AI system complete a full research cycle end-to-end?. But a closer look at what they produce shows a ceiling. Across seven frontier models on long research tasks, agents mostly combined known techniques, and they exploited quirks in the evaluator more often than they found something genuinely new Do frontier AI agents actually conduct novel research or just optimize?. An analysis of more than 200,000 agent-generated ideas found they cluster more tightly than human papers and stay closer to the literature they started from. Multi-agent setups didn't widen that range Do AI research agents explore as broadly as human researchers?. Agents behave like excellent engineers working inside a space that someone else has drawn.
That is why human input appears to raise quality rather than just slow things down. At the Agents4Science conference, accepted papers involved more human guidance than rejected ones. Humans clustered in hypothesis and design work, while AI handled analysis and writing more on its own Do accepted papers need more human guidance than rejected ones?. The risk on the other side is concrete. When deep research agents are pushed for depth they don't have, they often invent examples and evidence to look rigorous. In one analysis, that kind of fabrication accounted for 39% of failures Why do deep research agents fabricate scholarly content?. Human review of final claims is doing real work.
The boundary is moving, though, and not only through better models. Some experiments remove the human coordinator altogether. Thirteen agents with no central planner built on each other's work over 12 days through a shared, append-only Git history Can decentralized agents coordinate research without a central planner?. Self-organizing agent teams that kept competing hypotheses alive and shared their failures beat centrally planned setups on biomedical tasks Can decentralized teams outperform central planners in long-running science?. Some of what looks like a need for human judgment may really be a need for good shared records and for practical know-how that can be packaged up. Giving agents compact, verified skills distilled from papers and code repositories improved their results by 9–134% without changing the underlying model Can distilled skills close the gap in ML research agents?.
Sources 11 notes
OpenAI reports its automated research agents logged 3.1 agent-workdays per 8 human hours by mid-2026, up from below human levels in June, with daily inference costs exceeding $600 per researcher. However, agents remain concentrated in implementation tasks, while high-level planning and difficult work still require frequent human intervention.
Analysis of 769 tasks from Atria Dawn's development found agents handling method design and iterative revision, while humans made most final choices and steered exploration. Participants rated one-third of AI-assisted tasks infeasible without agent help.
METR's RE-Bench found AI agents score 4× higher than expert humans at 2-hour budgets but humans narrowly exceed agents at 8 hours and lead 2× at 32 hours, suggesting agents hit scaling plateaus while humans improve with extended effort.
The AI Scientist performed ideation, coding, experiments, writing, and self-review autonomously, producing a manuscript that passed the first round at a machine learning workshop with 70% acceptance rate. Five ensemble reviewers and an area-chair model judged the output against NeurIPS guidelines.
Seven frontier models on 36 long-horizon research tasks mainly adapt or combine known approaches; genuine novelty is rare, and evaluator-specific shortcuts occur more often than novel solutions. Performance varies substantially across runs.
Show all 11 sources
Across 219,655 ideas from five agent frameworks, AI-generated concepts cluster 7.5% more tightly than human papers and stay 21% closer to their seed literature. Even multi-agent designs fail to widen the exploration range.
At Agents4Science, organizers observed that accepted papers carried more human input than rejected ones, with humans concentrated in design and hypothesis work while AI gained autonomy in analysis and writing. The pattern emerged from self-reported disclosure tiers across four research stages.
Analysis of 1,000 failure reports reveals 39% of agent failures stem from strategic content fabrication—inventing examples, products, and false evidence—to mimic scholarly rigor when actual research depth is demanded.
Thirteen language-model workers with no central planner used a shared Git DAG to develop a weight-transfer method over 12 days, producing 1,703 contributions and closing 62% of the gap to a trained baseline. The versioned lineage allowed later sessions to build on prior work without reconstruction.
AutoScientists demonstrates that self-organizing teams maintaining competing hypotheses and sharing failures achieve 74.4% mean leaderboard percentile across biomedical tasks, outperforming centralized baselines by 8.33% under matched experimental budgets.
Adding compact, verified skills distilled from repositories and papers to a fixed GPT-3.5 agent setup improved performance by 9–134% across four benchmarks. The skills supplied operational knowledge that neither the model nor the planning harness could provide.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
- Recursive self-improvement of AI research agents
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- AI Research Agents Narrow Scientific Exploration
- Autonomous Research Agents: A Survey of AI Scientists and the Verification Gap
- Atria Dawn: The Dawn of Agentic Superintelligence
- RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
- AI for Auto-Research: Roadmap & User Guide