Autonomous Research Agents: A Survey of AI Scientists and the Verification Gap
Abstract Large language model (LLM) agents are increasingly deployed across the scientific research lifecycle: generating ideas, reviewing literature, designing and running experiments, analyzing results, drafting manuscripts, and reviewing them. End-to-end “AI scientist” systems now produce paper-like manuscripts that are evaluated by automated or workshop-style review. We argue that, in the public, full-text-coded sample we study, code release is now more common than reproducibility-grade and claim-verification artifacts: the open question is not just whether an agent can finish a research task, but whether anyone can verify the claims it produces. This survey makes that verification gap its organizing concern. We scope the survey to the computational (AI/ML) research setting, the testbed where autonomous-science claims are most observable because code, experiments, benchmarks, and write-ups can be inspected. We contribute four artifacts. First, a coded corpus screened from 125 candidates to 35 included works, of which we full-text code 26 entries (24 runnable systems plus two study/position works) on seven audit dimensions: lifecycle stage, autonomy level, evaluation method, released artifacts, human-in-the-loop points, novelty-verification method, and result-selection disclosure. The corpus yields the survey’s main pattern: code release is now common (83% of the 24 runnable systems), but the artifacts and checks that let a reviewer verify a result are not. Only 38% release the seeds or execution traces needed to reproduce a run, only 38% report any noveltyverification method, and the same qualitative pattern holds in the 22-system LLM-era runnable subset. These interpretive rates should be read as directional audit evidence because secondcoder agreement was lower for autonomy, novelty, and selection than for artifact release.
Introduction. Large language model (LLM) agents now attempt the full arc of research, from selecting a problem to writing the manuscript that reports the result [Lu et al., 2024, Yamada et al., 2025]. Systems described as “AI scientists” chain ideation, code generation, experiment execution, analysis, and manuscript writing into a single loop, and some report machine-generated manuscripts that are judged by an automated reviewer or at workshop level [Yamada et al., 2025, Miyai et al., 2025]. As such systems multiply, the question that matters changes. It is no longer whether an agent can complete a research task but whether we can trust the result, and whether anyone can check it at all [Bisht et al., 2026]. A system that emits a manuscript has not necessarily made a discovery: the claim may rest on a weak baseline, an unreproducible run, a hallucinated citation, or a result selected post hoc from many attempts. Recent critical work argues that today’s agents function as capable co-scientists yet are not built for autonomous discovery, citing biased problem selection, missing tacit laboratory knowledge, diversity collapse from preference optimization, and benchmarks that measure single-turn accuracy rather than closed-loop validity [Bisht et al., 2026]. Capability and verifiability have to be assessed together. We therefore organize the survey around verification rather than capability. The question is: what can autonomous research agents do without human judgment, and how would a reviewer know?
Why another survey. The area already has strong surveys, and we do not claim to be first. Zheng et al. [2025a] organize the field by autonomy level, through a Tool / Analyst / Scientist taxonomy viewed via the scientific method. Wei et al. [2025b] organize it by scientific domain, surveying autonomous discovery across the life sciences, chemistry, materials, and physics and unifying process, autonomy, and mechanism perspectives. The closest topical overlap is Tie et al. [2026], who directly survey “AI scientists” along a capability/workflow axis; recent deep-research surveys such as Xu and Peng [2025] catalogue retrieval-heavy research systems and their applications; and a separate line surveys scientific domain models such as molecular and protein LLMs [Zhang et al., 2024g], which are models, not agents, and we treat as context. Relative to all of these, our unit of analysis is the audit artifact a reviewer can inspect (released code, seeds, traces, selection policy, novelty-check method, human-intervention points), not the system’s domain or its autonomy label. Using audit artifacts as the unit of analysis lets us report quantitative disclosure rates and convert the survey into a reviewer checklist, rather than another descriptive taxonomy. Our scope is the AI/ML-research setting where those artifacts are inspectable. Table 1 summarizes the difference with graded, not binary, coverage.
• A coded corpus of autonomous research systems (Table 3), annotated on seven audit dimensions via a disclosed search protocol (Sec. 3); the corpus is the artifact every later claim rests on.
• A lifecycle × autonomy map (Table 6) indexed to the corpus, in which missing disclosures are coded explicitly rather than inferred.
• An auditability gap analysis (Table 9) linking evaluation proxies to recurring failure modes and the audit artifact each one needs.
How to read this survey. Figure 1 gives the historical orientation, and Figure 2 gives the paper’s internal map. We begin with a disclosed, full-text-coded corpus of autonomous research systems (Sec. 3, Table 3; full table in App. B). We then project those systems onto a lifecycle × autonomy map (Sec. 4, Table 6) to show where automation is claimed and where disclosure is missing. The map leads to the auditability gap analysis (Sec. 11, Table 9), which ties each evaluation proxy to the failure mode it may miss and the evidence that would make the claim easier to check. The verification ladder (Fig. 3) then separates domains with strong external checks from domains where the model’s own judgment often remains the main signal. The body of the paper walks through four lifecycle clusters: ideation and hypothesis (Sec. 5); literature and writing (Sec. 6); coding, execution, and analysis (Sec. 7); and review and closed-loop research (Sec. 8). Later sections reread the same systems by verification signal (Sec. 9) and measurement infrastructure (Sec. 10), before ending with a reviewer-facing checklist (Sec. 14, Table 10). Readers who want the shortest path can read Sections 3–4, the auditability gap analysis, and the checklist; the appendices hold the extended corpus, stage-coverage tables, verification-signal taxonomy, and benchmark catalogue.
Related work. Closest related surveys and verification-focused work. Recent work already argues that scientific agents need stronger verification, so our contribution is not the claim that verification matters. Gridach et al. [2025] survey agentic AI for scientific discovery across domains and emphasize evaluation, safety, and practical deployment challenges. Cornelio et al. [2025] make the verification problem explicit for AI-driven discovery, and Bisht et al. [2026] argue that current agentic AI scientists are not designed for autonomous discovery. Benchmark-disclosure audits show that agent evaluations often omit harness, cost, and reproducibility details [Moghadasi and Ghaderi, 2026], while chain-of-evidence systems such as ScientistOne [Meng et al., 2026] move toward explicit evidentiary traces. Our narrower contribution is to code public autonomous-research systems per system for audit artifacts, map those codings onto the research lifecycle, and convert the result into reviewer-facing reporting requirements. Prior work surveys agents, domains, workflows, or the need for verification; this paper asks what a reviewer can actually inspect for each system in the public record.
Method. Autonomous research did not begin with LLMs, and its predecessors make the verification contrast clear. We sketch three lineages the modern systems inherit from, and one they break from.
Pre-LLM lineages with built-in verification. The “Robot Scientist” program closed the hypothesis–experiment–interpretation loop for yeast genomics two decades ago [King et al., 2004], and its successor Adam became the first machine to autonomously discover novel scientific knowledge [King et al., 2009], with Eve extending the approach to drug repositioning [Williams et al., 2015]. This lineage logged hypotheses and provenance in machine-readable form, so each claim was auditable by construction. The AutoML tradition is similar: Auto-WEKA’s combined algorithmselection-and-hyperparameter problem [Thornton et al., 2013], Auto-sklearn’s meta-learning [Feurer et al., 2015], and the neural-architecture-search line from RL-based search [Zoph and Le, 2017] through transferable spaces [Zoph et al., 2018] to differentiable search [Liu et al., 2019] all operate over constrained search spaces with objective, automatically checkable success criteria. In these lineages, the search objective and the checking mechanism were usually coupled.
The agent substrate LLM research agents are built on. Modern research agents inherit a general toolkit: ReAct interleaves reasoning with actions [Yao et al., 2023b], Toolformer learns tool invocation [Schick et al., 2023], Tree-of-Thoughts turns reasoning into search [Yao et al., 2023a], Reflexion adds verbal self-correction [Shinn et al., 2023], and Voyager pursues open-ended skill acquisition [Wang et al., 2023a]. This machinery improves exploration, but self-critique, search, and memory do not by themselves guarantee that a claim, citation, or experiment has been independently checked. The capability transferred; the verification did not.
Where verification is built in, claims become easier to audit. The clearest audited LLMera successes pair generation with an automatic verifier. FunSearch gates LLM-proposed programs behind an executable evaluator and yields real mathematical advances [Romera-Paredes et al., 2024]; AlphaTensor accepts only algebraically valid tensor decompositions [Fawzi et al., 2022]; AlphaEvolve ranks evolved programs through task-specific tests [Novikov et al., 2025]; and Eureka keeps LLM-written reward code only when downstream training measurably improves [Ma et al., 2023]. The purest case is formal theorem proving, where a proof assistant (Lean, Coq, Isabelle) is a sound, deterministic verifier: an agent’s output is machine-checkable by construction, not judged by an LLM critic. This enables verifier-in-the-loop training and search (the DeepSeek- Prover line [Xin et al., 2024a,b, Ren et al., 2025], Kimina-Prover [Wang et al., 2025c], Goedel- Prover [Lin et al., 2025b]), retrieval-augmented Lean environments [Yang et al., 2023b], lifelong proof agents [Kumarappan et al., 2024], self-play conjecture-and-prove loops that turn the verifier into an open-ended discovery engine [Dong and Ma, 2025], and neuro-symbolic provers reaching medalist-level olympiad performance [Chervonyi et al., 2025]. These are not counterexamples to the verification framing but evidence for it: discovery is trustworthy precisely when a verifier, not the model’s own judgment, decides what counts. The rest of the survey asks what plays the verifier’s role when the domain admits no such checker.
Self-driving labs: verification by physical contact. Self-driving laboratories couple LLM reasoning to robotic execution, so claims must survive the physical world. A mobile robotic chemist searched a ten-dimensional photocatalyst space autonomously [Burger et al., 2020]; Coscientist designs and runs reactions via tool use [Boiko et al., 2023b]; ChemCrow augments an LLM with expert tools, and its authors report that GPT-4 acting as its own evaluator could not reliably separate correct from flawed outcomes [Bran et al., 2023]; A-Lab synthesized dozens of inorganic compounds, though later scrutiny of its phase-identification claims showed how “success” hinges on trustworthy characterization [Szymanski et al., 2023].
Discussion. The auditability gap links two questions prior surveys often treat separately: what current evaluation measures, and how these systems fail. A failure mode matters in proportion to how hard it is for a reviewer to detect.
What evaluation measures versus what trust requires. Most reported evaluation reduces to task success or a reviewer score. Trustworthy autonomous research instead requires evidence along dimensions that are rarely measured: scientific validity, novelty, reproducibility, experimental rigor, epistemic calibration, safety, and the true cost in compute and hidden human labor. The gap between these lists is the survey’s core observation: reported evaluations mostly measure task completion and infer scientific value. Novelty is the sharpest case, analyzed in detail in Section 5: while 38% of systems report some novelty-checking step, we found no system that reports independent validation that its novelty check is itself reliable.
Reading the gap table. Table 9 turns a list of complaints into a specification. Each row says: here is a way an autonomous-research claim can be wrong, here is why current evaluation does not catch it, and here is a minimum artifact or disclosure that would make the failure mode auditable. Position work arguing that current agents are not built for autonomous discovery [Bisht et al., 2026] and risk reports accompanying released systems [Miyai et al., 2025] supply much of the documented evidence; the rows marked “risk” are plausible modes we flag for systematic measurement rather than assert as established.
Autonomous research agents raise risks that scale with their autonomy. Here we treat them as auditability failures rather than as a complete safety taxonomy. Each maps onto a verification failure, which is why the reporting checklist that follows is also the survey’s safety instrument. Dualuse and biosecurity risk arises when an agent’s capability is not gated by red-team or task disclosure; it is now benchmarked directly (WMDP [Li et al., 2024c], SciSafeEval [Li et al., 2024d], the agentic bio-capabilities benchmark ABC-Bench [Liu et al., 2026a], and CBRN-risk quantification [Kumar et al., 2025]), with broader agent-safety [Zhang et al., 2024i] and controllable-risk frameworks for scientific agents [He et al., 2023b]. Research integrity fails when a claim’s soundness is not independently checked; new benchmarks test whether an agent can tell sound from unsound research (SoundnessBench [Ho et al., 2026]) and uphold academic integrity (SciIntegrity-Bench [Yang et al., 2026e]). Review-channel attacks exploit the absence of reviewer-independence and prompt-injection screening: LLM-written reviews are now detectable [Demetrio et al., 2025, Rao et al., 2025, Yu et al., 2025], manuscripts carry hidden prompts that hijack AI review [Lin, 2025, Collu et al., 2025], and tool-augmented detectors are emerging [Duarte et al., 2026, Duan and Li, 2026]. In each case the harm becomes possible exactly where an autonomous claim or review cannot be independently audited.
The audit problem grows fastest where agents add state, reward, self-modification, and oversight machinery. Continual memory, agentic reinforcement learning, self-evolution, and multi-stage pipelines raise the throughput of generated claims while adding surfaces that are hard to inspect. For each area, we ask what the mechanism automates, what verifies it today, and what a reviewer still cannot audit. The section closes by setting up the six open problems that follow.
Autonomy levels depend on the strength of the verification signal available. The literature divides along a mechanistic question: does the proposal try to expose an existing verification signal, or to manufacture one where none is available?
Conclusion. In the focal autonomous-research corpus, code release is more common than claim-verification evidence. The lifecycle × autonomy map shows competent stage-local and pipeline systems and an L4 column populated almost entirely by mechanical loops. Among the LLM-era systems, none shows an externally validated in-loop oracle under our coding rule; the one validated case, CAMEO, predates LLM agents and is included as a contrast benchmark. The auditability analysis then explains why task completion is not enough: validity, novelty, reproducibility, and selection bias require different evidence than a benchmark score supplies. The verification-signal taxonomy (Table 8) frames the agenda as movement from the model’s own judgment toward external checks. Six open problems follow. Proxy-to-truth validation for closed loops: most L4 systems optimize an internal metric, so the open question is when that proxy tracks scientific validity and how to detect when it does not. Closed-loop verifiers: evaluators that test whether results actually revise hypotheses, not merely whether a pipeline re-ran. Novelty auditing at literature scale: methods that verify a claim against the literature at the rate agents generate ideas, without treating LLM self-judgment as sufficient. Independent agent review: reviewers that are separate from the generator and robust to prompt injection and self-preference. Contamination-resistant evaluation: benchmarks whose validity survives train–test overlap and search-time leakage. Preregistration for AI-generated hypotheses: a shared record that curbs result selection and metric-driven problem selection [Bisht et al., 2026]. Until these are routine, an “AI scientist” that writes a manuscript should be read as automating the production of research artifacts, not yet providing the independent verification that would make its scientific claims trustworthy. The reporting checklist (Table 10) is a first step toward making that distinction auditable.
Limitations. Three limitations bound our claims. First, coding subjectivity and the reliability check’s scope: an independent second coder re-coded a random sample of ten systems on the four most interpretive dimensions, finding high agreement on artifact release (90%) but only 50–65% on autonomy level, novelty method, and selection disclosure. That check was performed from abstracts only, so it bounds the labeling subjectivity of those dimensions but does not independently validate the fulltext reading the headline rates rest on; a full-text second pass is the obvious next step. The one dimension with both high agreement and major interpretive weight is artifact release (90% agreement), which anchors the “code is common” half of our finding.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can AI systems perform peer review as effectively as humans?- Should AI research papers require dedicated automated review systems instead of existing journals?
- Can workshop acceptance rates reliably measure AI research quality compared to main conferences?
- Could AI feedback work as a substitute for human peer review entirely?
- What role should humans play in reviewing and approving AI-generated research?
- Do humans or AI perform better at different research stages?
- How does data availability shape which scientific questions AI systems tackle?
- Where does AI assistance become reliable versus prone to failure in science?
- How do template requirements limit AI research systems from true autonomy?
- What role should human experts play in AI-driven research ideation loops?
- What citation mistakes appear in fully autonomous AI research pipelines?
- How do researchers currently check whether an autonomous system's novelty claims are actually valid?
- What would a practical reviewer checklist for autonomous research systems need to include?
- Can AI agents themselves become reliable reviewers of other autonomous research systems?
- Can AI agents align their ideas with future research directions as well as humans do?
- Can human-AI collaboration preserve scientific breadth while improving individual productivity?
- How do autonomous science systems preserve competing hypotheses without a central planner?
- How do multi-agent writing systems maintain consistency across scientific manuscript sections?
- Does AI research acceleration compound into faster field-wide progress over time?
- Can recursive feedback loops turn AI research automation into genuine progress?
- What domains allow autonomous AI discovery because verification is fast enough?
- How does feedback latency from physical experiments shape AI system autonomy in research?