Towards End-to-End Automation of AI Research

Paper · arXiv 2606.15497 · Published March 31, 2026
Domain Specialization in LLMs

The automation of science is a long-standing ambition in the field of AI (Buchanan and Feigenbaum, 1978; Lenat, 1977). While the community has made significant progress in automating individual components of the scientific process, a system that autonomously navigates the entire research lifecycle—from conception to publication—has remained out of reach. Here, we present the strongest demonstration to date toward automating the entire process end-to-end. We present The AI Scientist, which creates research ideas, writes code, runs experiments, plots and analyzes data, writes the entire scientific manuscript and performs its own peer review. Its ideas, execution, and presentation are of sufficient quality to produce a manuscript generated by an AI system that passes the first round of peer review at a major machine learning conference workshop. The workshop has an acceptance rate of 70 percent. Our system leverages modern foundation models (Anthropic, 2024; Llama Team, 2024; OpenAI, 2023) within a complex agentic system. We evaluate The AI Scientist in two settings: a focused mode using human-provided code templates as an initial scaffold to conduct research on a specific topic, and a template-free, open-ended mode that leverages agentic search for wider scientific exploration (Chan et al., 2025; Jiang et al., 2025). Both settings produce diverse ideas and automatically test, report on, and evaluate them. This achievement demonstrates AI’s growing capacity for scientific contribution and signifies a potential paradigm shift in how research is conducted. As with any impactful new technology, there could be significant risks, including taxing overwhelmed review systems and adding noise to scientific literature. However, if developed responsibly, such autonomous systems could greatly accelerate scientific discovery.

Introduction. AI has long been used to aid scientific discovery, an ambition with deep roots in the history of the field (Langley, 2024; Langley et al., 1987; Lenat, 1977; Lenat and Brown, 1984; Waltz and Buchanan, 2009). Prior to the rise of large language models (LLMs), AI was limited to helping with specific, narrow tasks, such as discovering chemical structures (Buchanan and Feigenbaum, 1978), finding mathematical proofs (Lenat, 1977), discovering novel materials (Merchant et al., 2023; Pyzer-Knapp et al., 2022; Szymanski et al., 2023), and predicting the 3D shape of proteins (Hayes et al., 2025; Jumper et al., 2021). Other systems focused on analyzing pre-collected datasets to find novel insights (Falkenhainer and Michalski, 1986; Ifargan et al., 2025; Langley et al., 1987). However, with the recent advent of powerful and general foundation models, AI’s role has expanded to assist with a wider array of research activities. For example, LLMs now help with generating novel hypotheses (Faldor et al., 2024; Girotra et al., 2023; Hu et al., 2025; Lehman et al., 2023; Lu et al., 2025), writing literature reviews (Baek et al., 2025; Wang et al., 2024b), and coding experiments (Huang et al., 2024; Lu et al., 2024a; Ma et al., 2023; Zhang et al., 2025). Despite these advances in automating individual components, a system that autonomously navigates the entire research lifecycle—from conception to publication—has remained out of reach until now.

Method. The AI Scientist sequentially completes four main phases (Figure 1A): In the first phase, The AI Scientist is prompted to iteratively grow an archive (Mouret and Clune, 2015) of high-level research directions and hypotheses it can explore within a user-specified machine learning research subfield (an example progression is visualized in Supplementary Section C.4). For each direction, it generates a descriptive title, explanation of its reasoning for what the idea is and why it’s interesting to pursue it, and a proposed experimental plan (Supplementary Sections A.1.1 and A.2.6). After idea generation, The AI Scientist filters ideas by connecting the language model to the Semantic Scholar API (Fricke, 2018) and web access as tools (Schick et al., 2024). This allows The AI Scientist to discard any idea that is too similar to existing literature. The second phase of The AI Scientist executes the proposed experiments and then visualizes their results for the downstream write-up. We tested two different variants of experiment execution: (1) Template-Based: The AI Scientist is provided with a starting code template that reproduces a training run from a popular algorithm. The AI Scientist then executes the proposed experiment plan in linear order (Supplementary Section A.1). (2) Template-Free: Alternatively, The AI Scientist can generate an initial starting code script by itself. In this case, experimentation includes additional stages for optimizing the code it writes from scratch, and experiment execution leverages additional test-time compute with tree search (see methods). After each experiment, The AI Scientist is given the results and is prompted to take notes in the style of an experimental journal for future planning and writeup. The third phase of The AI Scientist produces a concise write-up of its research in the style of a standard machine learning conference paper. The AI Scientist is prompted to fill in a blank LaTeX conference template section by section, using its notes and plots (see methods). To construct the related work section and add citations throughout the manuscript, the system queries the Semantic Scholar (Fricke, 2018) API for relevant literature, comparing its findings against the generated manuscript over 20 rounds. For each potential citation, the system generates a textual justification for its inclusion, which informs The AI Scientist on how to use the reference appropriately within the manuscript. Finally, the paper generated by The AI Scientist undergoes a review by the Automated Reviewer to automatically evaluate the scientific quality of the conducted research.

The Automated Reviewer provides reviews based on the top-tier Neural Information Processing Systems (NeurIPS) conference review guidelines (NeurIPS Program Chairs, 2022). The output contains numerical scores (soundness, presentation, contribution, overall, and reviewer confidence), lists of weaknesses and strengths, as well as a binary decision (accept or reject). The Automated Reviewer’s pipeline consists of an ensemble of five reviews, followed by a meta-review where the model acts as an Area Chair to make a final decision conditioned on all five reviews (Supplementary Section A.3). We compared Automated Reviewer decisions with ground truth data for ICLR papers, extracted from the publicly available OpenReview dataset (González-Márquez and Kobak, 2024). As shown in Table 1, the Automated Reviewer’s agreement with human paper assessments is comparable to inter-human agreement measured by F1 and balanced accuracy as reported in the NeurIPS 2021 consistency study (Beygelzimer et al., 2021), which measured agreement between human reviewers on a comparable set of submissions (Supplementary Section A.3). This demonstrates its ability to replicate the collective judgment of human reviewers with high fidelity. These results are statistically significant (non-parametric bootstrap test (Efron and Tibshirani, 1993) and two-sample z-test (Lehmann, 1959); Supplementary Section A.3). Next, to investigate the effect of potential

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Can AI systems perform peer review as effectively as humans? What human oversight must AI research systems have? Does AI-assisted research sacrifice exploration breadth for productivity gains? Can AI research automation sustain progress through accelerating feedback loops? Should governance of agentic AI systems be runtime or design-time? How do educators verify student capability when AI can produce indistinguishable work? Can AI systems discover fundamental improvements to their own architectures?