Agentic AI Scientists Are Not Built For Autonomous Scientific Discovery

Paper · arXiv 2605.08956 · Published May 9, 2026
Domain Specialization in LLMs

A growing body of work pursues AI scientists capable of end-to-end autonomous scientific discovery. This position paper argues that although they already function as co-scientists, agentic AI scientists are not built for autonomous scientific discovery. We identify the following challenges in building and deploying autonomous AI scientists: (1) Problem selection is influenced by the McNamara fallacy; (2) Agents are built on large language models (LLMs) whose training corpora omit tacit procedural and failure knowledge of laboratory practice; (3) Preference optimisation during post-training compresses output diversity toward consensus; and (4) Most scientific benchmarks measure single-turn prediction accuracy and lack feedback from physical experiments back to the computational model. These challenges are not just questions of scale and scaffolding; they require revisiting fundamental design choices. To build truly autonomous AI scientists, we recommend scientific simulations as verifiers for training, a persistent, mutable epistemic state that carries beliefs and shifting objectives across an investigation, the establishment of a centralized preregistration repository for all AI-generated hypotheses, and application driven by scientific need rather than tool affordance.

Introduction. “In general, we look for a new law by the following process. First, we guess it; then we compute the consequences of the guess and then we compare the result of the computation to nature, with experiment or experience. If it disagrees with experiment, it is wrong. In that simple statement is the key to science.” — Richard P. Feynman (1964) One of the ways science is defined is the interrogation of nature through a cycle of conjecture, prediction, and physical verification. What distinguishes scientific reasoning from other forms of structured reasoning is precisely this closing of the loop against experiment—the willingness to be proven wrong by nature. Artificial intelligence (AI) has begun to participate meaningfully in this enterprise [Wang et al., 2023, Novikov et al., 2025]. Across the natural sciences, AI systems have accelerated the processing of large experimental datasets [Dagdelen et al., 2024], made progress towards automated tool-use and laboratories [Darvish et al., 2025, Mandal et al., 2025], uncovered latent structure in complex biological and chemical spaces [Jumper et al., 2021], and reduced the time from hypothesis to computational candidate in domains ranging from drug discovery [Coley et al., 2020a,b] to materials design [Zeni et al., 2025, Miret and Krishnan, 2025, Jain et al., 2013, Ahlawat et al., 2026].

Against this backdrop, a stronger proposition has gained traction: that large language model (LLM) based agentic systems can conduct natural science autonomously. A growing body of systems now claims end-to-end autonomous discovery—execution of a complete discovery cycle without meaningful human involvement: selecting problems worth pursuing, forming and revising hypotheses, designing and executing experiments, and communicating results [Kitano, 2021, Gottweis et al., 2025, Mitchener et al., 2025, Lu et al., 2026, Ghareeb et al., 2025, Yamada et al., 2025, Ghafarollahi and Buehler, 2025, Boiko et al., 2023]. We refer to such systems as AI scientists: LLMs deployed with the scaffolding of retrieval, tool use, critique, and memory. Major funding bodies are directing resources toward autonomous AI researchers [U.S. Department of Energy, 2026, Advanced Research + Invention Agency (ARIA), 2026]; benchmark performance on scientific tasks is being cited as evidence of near-human scientific capability; and the Nobel Turing Challenge frames the aspirational endpoint as an AI system capable of foundational discovery in the natural sciences [Kitano, 2021]. If these claims are correct, the organisation of scientific research faces genuine transformation. If they are not, then the mismatch between the framing and the actual capability carries costs for how systems are evaluated and scientific effort is allocated.

Here, we argue that current agentic systems are not built for autonomous natural science—not because current systems lack scale or tooling, but because the shortcomings are inherent to the training and deployment strategy. Unlike formal sciences such as mathematics where fast compilers and verifiers can address several of the shortcomings we identify, discovery in the natural sciences is bottlenecked by physical verification that is slow, partial, and itself requires interpretation [Bunge, 2017]; we centre our arguments there. We identify several challenges in building autonomous AI scientists, all the way from problem selection to application. The position advanced here is not that AI systems should be removed from scientific workflows—they are useful as a co-scientist: a capable collaborator performing one or more tasks in the discovery process, while at least one of problem specification, operation of laboratory instrumentation, physical execution of experiments, and interpretation of results remains with human scientists [Bianchi et al., 2026]. This is exemplified by AI systems’ success in problems that are well-stated by humans, primarily requiring optimisation in a well-defined problem space, and where constraints are defined a priori. Further, we do not claim that AI scientists are impossible, but argue that significant changes (Section 8) beyond scale and tooling are needed before they are feasible. The key contributions of the paper are as follows:

• We identify challenges along four themes (detailed in Sections 3– 6) that preclude autonomous scientific discovery: the McNamara fallacy in problem selection, the absence of tacit and failure knowledge from training corpora, diversity compression induced by preference optimisation, and the invalidity of current scientific benchmarks. • We provide empirical evidence of cross-model convergence, consistent with diversity compression, through the hypothesis hivemind experiment, showing that 12 frontier models from four providers, including open-weight models, converge semantically on both interpretive and open-ended hypothesis generation tasks in two natural-science domains.

Related work. The aspiration to automate scientific discovery predates modern AI by several decades [Lindsay et al., 1993, Langley, 1977, King et al., 2009] . LLM-based agentic systems have renewed this ambition at a different scale. AI systems have proposed a drug candidate for dry age-related macular degeneration [Ghareeb et al., 2025], produced scientific reports spanning multiple domains in a single turn [Mitchener et al., 2025, Ghafarollahi and Buehler, 2025], and generated millions of candidate crystal structures, compressing what would have been decades of screening into days [Merchant et al., 2023, Zeni et al., 2025]. However, for each of these examples, human researchers still specified the problem, provided the dataset, and defined the criteria for success. We argue that this division of labour is not merely a short-term limitation fixable with scale but follows from design choices that do not fully reflect the spirit of scientific inquiry. The challenges we identify scale with the cost and completeness of the signal that tells a system it was wrong. In mathematics a proof assistant either verifies a result or returns a detailed error trace within seconds, and autonomous discovery is correspondingly more tractable; we do not dispute claims of autonomy in such domains. In materials discovery a density functional theory calculation takes hours and returns only a proxy for stability, while a synthesis attempt can take months and yields a result that must itself be interpreted before it counts as evidence.

Current systems that claim to be AI scientists are composed of a small and recurring set of components: retrieval over published literature; execution of tools and generated code; decomposition into multiple agents holding distinct roles, usually including a critic; iterative refinement of a candidate against that critic; memory or a shared notebook carried between calls; and, in a minority of systems, actuation of laboratory hardware [Boiko et al., 2023, Darvish et al., 2025, Ghareeb et al., 2025, Ghafarollahi and Buehler, 2025, Mitchener et al., 2025, Lu et al., 2026].

While this scaffolding improves these systems compared to a single LLM, the challenges we identify remain: (1) Problem selection happens before the scaffold is invoked; (2) Retrieval changes what the model can access but not what exists to be accessed, and thus cannot compensate for the tacit and failure knowledge gaps; (3) Critique and re-ranking by agents select among samples already drawn from the base model, reordering them without adding weight to regions that post-training has thinned (consistent with this, Ríos-García et al.

Method. To ground our arguments, we discuss the search for solid-state battery electrolytes (SSE) with high ionic conductivity and electrochemical stability as a running example. SSE is one of the most impactful and well-studied problems at the frontier of energy storage research [Dutra et al., 2025]. The atomistic mechanisms governing ion transport in solid electrolytes remain incompletely understood. No single readily computable objective captures all practical performance requirements. Progress demands coherent reasoning across scales—from quantum-mechanical density functional theory calculations, through molecular dynamics simulations capturing diffusive ion motion, to devicescale electrochemical models—each with its own intricacies. Finally, realizing the computationally designed electrolyte in a laboratory setting presents unique challenges including synthesizability, processability, and device-scale manufacturing. SSE discovery sits at the far end of the easy to hard verification spectrum described, and the complexity and urgency of the task make it a good backdrop to understand the challenges we outline.

Problem selection is categorically different from problem solving; it involves judgment that integrates theoretical significance, practical tractability, community need, resource constraints, and the innate human need to understand the universe in ways that resist formal specification [Nickles et al., 1980, Polanyi, 1966]. In Section 2 we stated that current AI scientists rely on human judgement for problem selection; here we extend the argument by showing that the presence of AI technologies negatively impacts the human judgement they rely on.

The McNamara or Quantitative fallacy describes what happens when institutions face complex problems full of intangibles: They measure what is easy to measure, disregard what cannot be quantified, presume the unmeasured is unimportant, and finally declare that the unmeasured does not even exist [Yankelovich, 1972, O’Mahony, 2017]. AI techniques rely on numerical signals such as losses, rewards, or fitness values for optimization and evaluation of the final system, which naturally pushes research projects involving AI toward problems that can be quantified. In addition, machine learning techniques often require large amounts of data in a format that’s easy to manipulate computationally, further reducing the span of problems amenable to AI solutions. This effect is already documented in scientific research. Across 41.3 million research papers, Hao et al. [2026] found that AI-augmented scientists publish 3.02 times more papers and receive 4.84 times more citations, yet AI adoption shrinks the collective volume of scientific topics studied by 4.63% and reduces scientist-to-scientist engagement by 22%. The AI tools accelerate established fields leading to a narrower set of questions being asked, while genuinely hard, data-sparse problems go untouched.

Scientific publishing filters for some kinds of scientific knowledge over others by design or bias in the peer-review process. It selects for conclusions over process, for positive results over negative ones, and for ideas consistent with prevailing theory over those that challenge it. A model trained on this corpus does not inherit a snapshot of scientific knowledge but the output of that filter. Two types of omission matter most, and neither can be recovered by training on more of the same data.

Tacit knowledge Laboratory practice produces understanding that is rarely written down: which synthesis conditions are reliable, which reagents behave inconsistently across suppliers, which reported protocols require undocumented adjustments, and which measurements carry signatures of known experimental artifacts. Polanyi [1966] argued that this kind of knowledge is irreducibly nonpropositional; an experienced researcher can recognise a result that looks too clean or a curve with the wrong shape, but cannot fully articulate the basis for that recognition that could be published or trained upon [Fjelland, 2020].

Discussion. 7 Alternative views Scale and scaffolding will close the autonomy gap. Improvements in context length, retrieval reliability, and tool-use accuracy will improve capabilities, but the failures described in this paper sit upstream of these properties (Section 2): more positive results cannot teach what failed experiments would, and a scaled-up model trained toward annotator consensus is likely to converge on the same region of hypothesis space at higher fluency. Experiments show that agentic scaffolding also fails to correct for this [Ríos-García et al., 2026].

Demonstrated successes like AlphaFold and GNoME show that full autonomy is within reach. On the contrary, these contributions exhibit the boundary the paper is trying to draw. AlphaFold solved a problem formally stated for fifty years, with an unambiguous evaluation criterion available at experimental scale. GNoME searched a compositional space against a computable stability criterion. In both cases, human researchers specified the objective, defined the validation criterion, and interpreted what the result meant for subsequent work. Extrapolating these successes to autonomy assumes that this human judgment is unnecessary or reproducible by the same paradigm, which the evidence does not support. Bianchi et al. [2026] found that the highest-quality scientific outputs in a controlled study came from meaningful human-AI collaboration, with quality declining under full automation. The successes of the current paradigm are an argument for investing in the co-scientist model, not for replacing the human half of it.

The distinction between co-scientist and autonomous scientist collapses as systems improve. If human oversight becomes increasingly nominal as capability grows, co-scientist might be a label applied to a workflow that is functionally autonomous. We argue that human involvement will not diminish randomly or uniformly; it will recede in precisely those aspects where the problems we identify are addressed, and these require different design commitments rather than more capability of the current kind. Until those are in place, the co-scientist model is an accurate description of what the collaboration requires: human judgment that supplements models where they are the most limited.

Human scientists are also biased, limited, and frequently wrong. Individual human scientists are subject to confirmation bias, motivated reasoning, and the same consensus pressures that preference optimisation encodes [Wason, 1960]. We believe the relevant comparison is between an AI system and the collective enterprise of science with corrective mechanisms such as peer review, replication requirements, adversarial collaboration, and reputational accountability for error developed over decades [Longino, 1990]. A researcher who is wrong about a hypothesis faces replication attempts and critical commentary; but the convergence documented in Section 5 means that querying additional AI systems is unlikely to apply the same corrective pressure.

8 Toward autonomous AI scientists Autonomous AI scientists nonetheless remain a goal worth pursuing. The limited availability of expert human judgement can constrain the rate of scientific progress. Further, long-horizon high-risk problems are attempted rarely because scientists are incentivised toward incremental projects [Foster et al., 2015]. The scarcity of human judgement and risk-aversion can be reduced with highly capable AI scientists. We discuss some steps to make progress towards autonomous discovery.

Conclusion.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

What human oversight must AI research systems have? Does AI-assisted research sacrifice exploration breadth for productivity gains? Can AI systems discover fundamental improvements to their own architectures? Can AI systems achieve real improvement without external human feedback? How do real-world evaluations reveal AI capabilities that benchmarks hide? Can mechanistic interpretability methods reliably reveal what models actually know? How should humans and AI agents share control and decision-making? Why does AI verification capability persistently exceed generation capability? Do individually safe AI actions create unsafe outcomes in integrated systems? How do AI systems determine and balance multiple competing objectives? What explains the gap between benchmark scores and true reasoning capability?