Can AI agents produce scientifically novel and important ideas?
A 2025 conference let AI agents lead research and peer review, accepting 48 of 314 papers. Reviewers found technically sound work but questioned whether it addressed questions that actually matter to science.
Science News reports that Agents4Science 2025, a virtual conference held Oct. 22, 2025, accepted papers from any field of science on one condition: AI agents had to do most of the work, from formulating hypotheses and analyzing data to the first round of peer review, with humans assessing only the top submissions. By the conference's count, 48 of 314 submissions were accepted, and each team had to detail how people and AI collaborated at every stage. Most of the evidence on quality comes from participants. Min Min Fong's team found AI "really great" for "computational acceleration," but it kept citing the wrong date for when San Francisco's towing-fee waiver took effect, and Fong caught the error only by checking the original source. Risa Wechsler, who helped review submissions, found the papers "technically correct" but "neither interesting nor important," and warned that technical skill can "mask poor scientific judgment."
Organizer James Zou notes that most journals and meetings ban AI coauthors and bar reviewers from relying on AI, to avoid hallucinations, and that this makes it hard to learn how good AI is at science. The conference answers by letting agents carry the early stages and placing human judgment at selection. The pattern the excerpt shows is a split: computation and first-pass work are delegated, while error-checking and judgment about which questions matter stay with people. One counterexample complicates it. Silvia Terragni, a machine learning engineer at Upwork, asked ChatGPT to propose paper ideas from her company's problems, and one result was among the three top papers. "I think [AI] can actually come up with novel ideas," she says.
The excerpt sits on the boundary that Where does AI assistance become unreliable in research? describes. Fong's computational help falls on the reliable side and Wechsler's critique on the unreliable side. Terragni's idea fits neither side cleanly, so the excerpt extends the stage boundary rather than confirming it: the same kind of system is credited with novelty in one case and faulted for unimportance in another. The conference also inverts the arrangement in How should AI agents and humans divide research tasks?, where humans keep most final decisions; here agents do the first-round review and humans judge only the top submissions. It is different evidence from Can specialized agents write better scientific papers than single models?: PaperOrchestra's figures come from a comparison its builders ran, while these verdicts come from outside expert reviewers.
The excerpt does not establish much about quality. It is one journalist's account of one event, with no per-paper scoring, no reviewer agreement figures, no error count across the 48 accepted papers, and no comparison with human-written submissions judged under the same process. The strongest positive claim, that AI can generate novel ideas, rests on one top-three paper described through its author's account. What the excerpt supports is narrower: agents can produce work that clears a human selection process at a venue that required disclosure, and the expert who judged it found it technically sound but missing the judgment that makes work matter. At that strength, autonomy in the early stages of research looks checkable mainly by people with domain expertise, so the verification step remains the bottleneck.
Inquiring lines that read this note 1
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Does AI-assisted research sacrifice exploration breadth for productivity gains?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Where does AI assistance become unreliable in research?
This explores whether AI capability follows a sharp boundary in research tasks, and what determines which side of that line a task falls on. Understanding this matters because it reveals where humans must stay in control.
this excerpt's computation-versus-judgment split tests that boundary, with a novelty case it does not explain
-
How should AI agents and humans divide research tasks?
In building its own foundation model, Atria Dawn studied how to split work between agents and human researchers. Understanding this division matters for designing effective human-AI collaboration in technical R&D.
the lab keeps humans on final decisions; the conference gives agents the first-round review instead
-
Can specialized agents write better scientific papers than single models?
Multi-agent frameworks decompose writing into specialized subtasks. This explores whether distributed agents maintaining cross-document consistency outperform single-model approaches on manuscript quality and literature synthesis.
a builder-run comparison; this conference reports outside expert verdicts, a different kind of evidence
-
Do accepted papers need more human guidance than rejected ones?
Agents4Science organizers reported that accepted papers involved more human input than rejected papers, with humans leading design and AI handling analysis. This raises whether human guidance predicts acceptance and how labor should divide in AI-authored research.
qualifies: accepted papers carried more human guidance, with humans shaping design and hypothesis while AI took later stages
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- A science conference tested AI agents as authors and reviewers
- Exploring the use of AI authors and reviewers at Agents4Science
- Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?
- AI Research Agents Narrow Scientific Exploration
- What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
- Autonomous Research Agents: A Survey of AI Scientists and the Verification Gap
- Recursive self-improvement of AI research agents
- Atria Dawn: The Dawn of Agentic Superintelligence
Original note title
Agents4Science 2025 accepted 48 of 314 AI-led papers — one reviewer found them technically correct but neither interesting nor important