Accelerating scientific discovery with Co-Scientist

Paper · arXiv 2502.18864 · Published February 26, 2025
Domain Specialization in LLMs

Summary Scientific discovery is driven by scientists generating novel hypotheses for complex problems that undergo rigorous experimental validation. To augment this process, we introduce Co-Scientist, a multi-agent AI system built on Gemini for structured scientific thinking and hypothesis generation. Co-Scientist aims to help scientists discover new original knowledge. Conditioned on their research objectives and prior scientific evidence, it formulates demonstrably novel research hypotheses for experimental verification. The system’s design involves agents continuously generating, critiquing and refining hypotheses accelerated by scaling test-time compute. Key contributions include: (1) a multi-agent architecture with an asynchronous task execution framework for flexible compute scaling; (2) a tournament evolution process for self-improving hypotheses generation. Automated evaluations show continued benefits of test-time compute scaling, improving hypothesis quality over time. While general purpose, we focus the validation in three biomedical applications: drug repurposing, novel target discovery 1, and explaining mechanisms of anti-microbial resistance 2. Specifically, Co-Scientist helped identify new drug repurposing candidates and synergistic combination therapies for acute myeloid leukemia, which were validated through in vitro experiments. These real-world validations demonstrate the potential of Co-Scientist to accelerate scientific discovery and usher in an era of AI empowered scientists. Introduction Researchers are faced with a breadth and depth conundrum. The complexity of scientific topics require increasingly deep and specific subject matter expertise, while leaps in insight may still arise from broad knowledge bridging across disciplines 3–5. With the rapid rise in scientific publications and the development of numerous specialized technologies, mastery of both discipline-specific depth and trans-disciplinary insight can be challenging.

Introduction. Co-Scientist works through a significant scaling of the test-time compute paradigm 16–18 implementing structured scientific thinking in a multi-agent setup to iteratively reason, evolve, and improve the outputs as it gathers more knowledge (Fig. 1b). Underpinning the system are thinking and reasoning steps—notably a self-play based scientific debate step for generating novel research hypotheses; tournaments that compare and rank hypotheses via the process of finding win and loss patterns, and an evolution process to improve their quality. Finally, the agentic nature of the system enables it to recursively self-critique its output and use tools such as web-search and specialized AI models to provide itself with feedback to refine its hypotheses and research proposals.

While Co-Scientist is general purpose and applicable across scientific disciplines, we validate it in three impactful areas of biomedicine with varied complexity: drug repurposing for cancer, novel treatment target discovery for liver fibrosis, and new mechanistic explanations for antimicrobial resistance (Fig. 1c).

Drug development remains an expensive and protracted process, with most new approvals requiring de novo discovery for each indication 19. Systematic identification of new therapeutic indications for approved agents via drug repurposing offers a pragmatic strategy to accelerate development timelines and reduce attrition 20. Using Co-Scientist, we generated large-scale repurposing predictions validated through expert curation and in vitro assays. The system proposed several single-agent and combination therapies for AML that demonstrate selective cytotoxicity at clinically relevant concentrations. Beyond repurposing, Co-Scientist enables hypothesis generation for de novo target discovery, a process traditionally limited by the scale and uncertainty of biological inference. We applied Co-Scientist to liver fibrosis, where it proposed and ranked novel epigenetic targets demonstrating significant anti-fibrotic activity and hepatocyte regeneration in human hepatic organoids 1. Finally, we examined bacterial gene transfer mechanisms related to antimicrobial resistance (AMR), a system-level challenge involving molecular mechanisms and evolutionary pressures 2. Researchers instructed the system to explore a topic their group had independently discovered, but not yet published. Co-Scientist was asked to hypothesize how capsid-forming phage-inducible chromosomal islands (cf-PICIs) exist across bacterial species. It independently proposed that cf-PICIs interact with diverse phage Overall, our key contributions are summarized as follows: (1) Introducing Co-Scientist. We develop and introduce Co-Scientist, a structured scientific thinking engine, that goes beyond literature summarization and “deep research” tools to assist scientists in uncovering new knowledge, novel hypothesis generation, discovering unexpected connections, and experimental planning. (2) Significant scaling of the test-time compute paradigm for scientific reasoning. Co-Scientist is built on a Gemini based multi-agent architecture, utilizing an asynchronous task execution framework. This framework allows the system to flexibly allocate computational resources to scientific reasoning, mirroring key aspects of the scientific method. Specifically, the system uses self-play strategies, including a scientific debate and a tournament-based evolution process, to iteratively refine hypotheses and research proposals creating a self-improving loop. Using automated evaluations across 15 complex expert curated open scientific goals, we demonstrate the benefits of scaling the test-time compute paradigm with Co-Scientist outperforming other state-of-the-art (SOTA) agentic and reasoning models in generating high quality hypotheses for complex problems. (3) Expert-in-the-loop scientific workflow. Our system is designed for collaboration with scientists. The system can flexibly incorporate conversational feedback in natural language from scientists and co-develop, evolve and refine outputs. (4) End-to-end validation of Co-Scientist in important topics in biomedicine. We present end-to-end validation of novel AI-generated hypotheses through new empirical findings in three distinct and increasingly complex areas of biomedicine: drug repurposing, novel target discovery, and antimicrobial resistance (Table 1, Supplementary Note 1). Table 1 | Three real-world applications in biomedicine for end-to-end validation of Co-Scientist. The table summarizes three scientific tasks selected to evaluate the hypothesis generation capabilities of Co-Scientist. The chosen applications span varying biological disciplines and are categorized by four axes, inherent challenge (the primary scientific objective), complexity (the depth of reasoning required), scale (data availability and experimental feasibility), and unknown elements (the boundaries of the hypothesis search space).

Method. Co-Scientist overview Given a research goal, Co-Scientist generates hypotheses constrained by default criteria including plausibility, novelty, testability, and safety. At a high level, it employs an asynchronous multi-agent architecture where a team of agents co-operate to solve scientific problems and develop novel hypotheses (Fig. 1b). It comprises a natural language interface for expert supervision, a task execution framework for resource allocation, a suite of specialized agents (Generation, Reflection, Ranking, Evolution, Proximity, Meta-review) mirroring the scientific method, and a persistent context memory for long-horizon reasoning. Detailed system configurations and agent mechanisms are fully described in the Methods. System analysis and evaluation We first conduct the initial system evaluations to benchmark and verify the choice of the architecture and metrics underpinning Co-Scientist (detailed in Supplementary Note 2, Supplementary Fig. 1). We perform an ablation study to investigate the contribution of each agentic component in Co-Scientist. We then analyze the impact of scaling test-time compute, and undertake a small-scale evaluation with domain experts to assess the quality of the system outputs. Finally, to assess the practical utility of the system’s novel predictions, we perform end-to-end wet-lab validations (laboratory experiments) of Co-Scientist-generated hypotheses and research proposals in three key biomedical applications: drug repurposing, discovering novel treatment targets, and elucidating the mechanisms underlying antimicrobial resistance (Table 1, Supplementary Note 1). The varying complexity and nature of these applications enable a more comprehensive assessment of the system. Notably, all three validations involved expert-in-the-loop. Agent ablation analysis. Our ablation analyses, detailed in Methods, Supplementary Note 3, and Supplementary Fig. 2-6, confirmed the importance of our multi-agent architecture and specialized prompting strategies for robust scientific reasoning. For instance, granting the Reflection agent access to external search tools effectively prevented the hallucination of seemingly novel but implausible hypotheses, while employing a scientific debate prompt in the Ranking agent significantly improved the ranking of hypotheses and reduced positional bias. Furthermore, iterative refinement by the Evolution agent substantially boosted hypotheses quality. Scaling test-time compute improves scientific reasoning. To evaluate the effects of test-time compute scaling and Co-Scientist’s progress during iterative scientific reasoning and hypothesis generation, we measured the Elo ratings of Co-Scientist generated hypotheses and proposals over the course of its thinking and computation (i.e. the tournament of hypotheses). This analysis was done across 203 distinct research goals curated across broad scientific topics (predominantly in biomedicine, but also included other topics such as mathematics and physics) and entered into Co-Scientist until February 3, 2025.

Discussion. In this work, we report the development and initial validation of a multi-agent, Gemini based AI system, Co-Scientist, designed as a structured scientific thinking engine to accelerate novel scientific discovery. Co-Scientist moves beyond conventional computational approaches through the in-silico implementation of a multi-agent architecture that mirrors the core aspects of the scientific method. Instead of brute-force generation, the system iteratively refines hypotheses through a “generate, debate, evolve” paradigm. This method, which incorporates self-debate, tournament-based selection, and iterative evolution and refinement, enables a progressive convergence on high-quality, well-supported hypotheses, thereby scaling research ideation with test-time compute rather than exhaustive generation. The system’s context memory, combined with the iterative self-improvement cycle, functions as an emergent internal model of the scientific research process. While not an explicit symbolic model, it represents a progressively more coherent and interconnected state of knowledge, facilitating the synthesis of information and the identification of knowledge gaps.

The practical utility of this approach was demonstrated through the generation of novel and experimentally tractable hypotheses across three challenging and varied biomedical problems. In oncology, Co-Scientist identified drug repurposing candidates for AML that showed in vitro efficacy at clinically relevant concentrations. For liver fibrosis, it proposed novel epigenetic targets, leading to the experimental validation of several anti-fibrotic compounds, including one FDA-approved drug. Furthermore, in microbiology, the system independently recapitulated a novel, (and then) unpublished mechanism of mobile genetic element transfer between bacteria. These findings provide preliminary evidence that Co-Scientist can contribute meaningfully to scientific discovery by amplifying scientists.

This system’s architecture is model-agnostic, allowing it to leverage the advancing capabilities of frontier LLMs without requiring retraining of the whole agentic framework such as Gemini 3, GPT 5.4 and Opus 4.6. With the latest advances in frontier models, we expect further significant improvements in the quality of hypotheses generated and the complexity of scientific tasks the system can autonomously accomplish.

Conclusion. The continued development of Co-Scientist will focus on three key areas. Immediate improvements will target the system’s robustness by enhancing learning and knowledge base, literature search capabilities to broaden access, implementing more rigorous fact-checking against external databases and tools, and improving citation recall. Future advancements will focus on expanding the system’s core capabilities. This includes integrating agents that can directly reason over public databases and multimodal data, enabling bioinformatics and data science tasks. The implementation of reinforcement learning from human and experimental feedback could further optimize the hypothesis generation and refinement process.

Expanded evaluations are also necessary to assess Co-Scientist’s generalizability across a wider range of scientific disciplines. This requires developing more objective and automated evaluation metrics that move beyond current ranking systems and engaging a larger cohort of domain experts to stress-test the system with diverse and complex research queries.

In the fullness of time, integrating Co-Scientist with laboratory automation platforms could create a closed-loop, autonomous system for hypothesis generation, experimental validation, and iterative learning, significantly accelerating the pace of scientific discovery. Conclusion Co-Scientist represents a promising step towards AI-assisted augmentation of scientists and acceleration of scientific discovery. Its ability to think scientifically, generate novel testable hypotheses across diverse scientific and biomedical domains, some supported by experimental findings, along with the capacity for recursive self-improvement with increasing compute, demonstrates the promise of meaningfully accelerating scientists’ endeavors to resolve grand challenges in human health, medicine and science. This innovation opens numerous questions and opportunities.

Limitations. Despite these promising early results, several limitations must be addressed. Co-Scientist’s knowledge is constrained by its reliance on open-access scientific literature, which may lead to the omission of critical prior art behind paywalls and a systemic lack of access to negative experimental results. Furthermore, the quality of generated hypotheses relies on the mixed and contradictory quality of the source literature; thus, there is a risk of propagating erroneous or irreproducible findings. A key future direction is the development of agents with enhanced provenance capabilities to trace claims to specific figures or data within a source, mitigating the impact of unreliable literature.

Co-Scientist also inherits the intrinsic limitations of its underlying models, including imperfect factuality and the potential for hallucinations. Improving reasoning capabilities is a critical area for future work. Additionally, the validation of Co-Scientist’s hypotheses, while successful, remains preliminary.

Finally, the broader integration of such AI systems into the scientific workflow requires careful consideration of potential bias, which could risk diminishing critical thinking or homogenizing research directions.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Does AI-assisted research sacrifice exploration breadth for productivity gains? What human oversight must AI research systems have? Can AI systems perform peer review as effectively as humans? Why do LLM research ideation systems generate novelty but lack diversity? Can AI systems discover fundamental improvements to their own architectures? How do clinicians calibrate trust in AI medical recommendations? Why does AI verification capability persistently exceed generation capability?