Artificial Intelligence Tools Expand Scientists' Impact but Contract Science's Focus (Just accepted by Nature, to be online soon)
Development in Artificial Intelligence (AI) has accelerated scientific discovery1. Alongside recent AI-oriented Nobel prizes2–9, these trends establish the role of AI tools in science10. This advancement raises questions about the potential influences of AI tools on scientists and science as a whole, and highlights a potential conflict between individual and collective benefits11. To evaluate, we used a pretrained language model to identify AI-augmented research, with an F1-score of 0.875 in validation against expert-labeled data. Using a dataset of 41.3 million research papers across natural science and covering distinct eras of AI, here we show an accelerated adoption of AI tools among scientists and consistent professional advantages associated with AI usage, but a collective narrowing of scientific focus. Scientists who engage in AI-augmented research publish 3.02 times more papers, receive 4.84 times more citations, and become research project leaders 1.37 years earlier than those who do not. By contrast, AI adoption shrinks the collective volume of scientific topics studied by 4.63% and decreases scientist’s engagement with one another by 22.00%. Thereby, AI adoption in science presents a seeming paradox—an expansion of individual scientists’ impact but a contraction in collective science’s reach—as AI-augmented work moves collectively toward areas richest in data. With reduced follow-on engagement, AI tools appear to automate established fields rather than explore new ones, highlighting a tension between personal advancement and collective scientific progress.
Introduction. Artificial intelligence (AI) has made significant strides in recent decades, promising to impact myriad aspects of society, including education12,13, healthcare14,15, and industry16. Major investments in predictive and generative AI have catalyzed society-level debates over the future of AI at home and in the workplace. Perhaps more than any other domain, AI tools have become deeply entwined with the process of knowledge production, yielding findings that attract disproportionate attention in various scientific fields1. For example, AlphaFold learns known protein structures to accurately predict the unexplored ones, circumventing the capital and human cost of conventional structural inference and recently granted a 2024 Nobel Prize9,17. Models improved via deep reinforcement learning have become tuned to contain complex fusion reactions18 and discovered new, hardware-optimized forms to matrix multiplication that recursively accelerate deep learning itself19. Autonomous laboratory systems driven by ChatGPT have helped some chemists and material scientists upscale the number of adaptive high-throughput experiments20–22. Moreover, recent developments in large language models are making them increasingly incorporated in assisting scientific writing23–26, facilitating the distillation of scientific findings, but also raising concerns about weakened confidence in AI-generated content21,22,27. AI’s increasing capabilities to influence scientific research suggest that it manifests potential to both increase the productivity of individual scientists and raise the visibility of science it supports. Despite the increasing adoption of AI in science, large-scale empirical measurements of AI’s scientific impact are limited, and a detailed, dynamic understanding of AI’s impact on the entire character of science remains largely unknown. Recent work suggests AI has brought widespread benefits to individual scientists but may lead to demographic disparity resulting from gaps in AI education10. Researchers have also identified evolving citation patterns that signal a changing scientific landscape in AI research28. Here we seek to explore the impact of AI in scientific research at different scales by posing the question: How does the adoption of AI influence individual scientists’ careers and the collective exploration of science as a whole? We conduct a large-scale quantitative analysis of the impact of AI on scientists and science, covering 41,298,433 research papers spanning from 1980 to 2025 in the OpenAlex dataset29, with patterns corroborated using the Web of Science30,31.
Notably, we do not focus on computer science or mathematics, fields that develop AI methodologies directly, but rather on papers that augment research in natural science fields by adopting AI, primarily covering decades involving development and deployment of conventional machine learning algorithms and also extending to a necessarily more preliminary analysis of the latest generative AI techniques. Specifically, we select six representative disciplines that cover the vast majority of natural science contributions—biology, medicine, chemistry, physics, materials science, and geology. We then leverage the BERT language model32,33 to accurately identify such AI-augmented research papers based on their titles and abstracts. We separate the periods in which AI was predominantly conventional machine learning, deep learning, and most recently generative designs including large language models. With abundant data-based evidence across decades of conventional machine learning and deep learning, we validate these AI-based measurements and use them to reveal that the adoption of AI leads to an amplifying effect on the career of individual scientists, bringing acceleration in the production and visibility of science produced by those scientists who incorporate AI. Nevertheless, this effect corresponds with a contracted focus within collective science. Measured with “knowledge extent”, the “diameter” covered by a sampled batch of papers in vector space, AI-driven science spans less topical ground and is associated with a decrease in follow-on scientific engagement, suggesting that AI is currently more likely to focus on existing popular research problems rather than explore new ones. Meanwhile, analyses using currently available data within the latest era of generative AI including large language models reveal a preliminary consistency with prior periods, providing a starting point for further study as generative AI develops over a longer period.
Method. M1. Dataset and Paper selection In this section, we introduce the procedure of selecting the research papers included in our analysis. In this paper, we conduct our major analyses based on OpenAlex29. OpenAlex is a scientific research database built upon the foundation of the Microsoft Academic Graph (MAG)52,53. Supported by non-profit organizations, OpenAlex is continuously updated, providing a sustainable global resource for research information. As of March 2025, OpenAlex contains 265.7M research papers, along with related data about citation, author, institution, etc. Among the massive quantity of papers in the OpenAlex dataset, we select 66,117,158 English research papers published in journals and conferences spanning from 1980 to 2025 and filter out those with incomplete titles or abstracts. We identify the scientific discipline each paper belongs to utilizing the topics contained in OpenAlex, which are extracted using a natural language processing approach that annotates titles and abstracts with Wikipedia article titles as topics sharing textual similarity. In the raw dataset, these topics form a hierarchical structure and each paper is associated with several. Adopting the 19 basic scientific disciplines in the Microsoft Academic Graph (MAG)52,53, i.e., art, biology, business, chemistry, computer science, economics, engineering, environmental science, geography, geology, history, materials science, mathematics, medicine, philosophy, physics, political science, psychology, and sociology, we trace along the hierarchy to determine to which disciplines each topic belongs. We note that because the original topics of one paper may be retraced to different topics, the scientific discipline of each paper may not be unique. In other words, one paper may span two or more academic disciplines, e.g., chemistry and biology, which reflects the common phenomena of borderline or interdisciplinary research54. In this paper, we emphasize the adoption of AI methods in conventional natural science disciplines and exclude research developing AI methodologies themselves, separating the influence of AI on science from AI’s own invention and refinement. Therefore, we select biology, medicine, chemistry, physics, materials science, and geology as representatives of natural science disciplines, while we exclude computer science and mathematics, where most works introducing and developing AI methods are published. We also exclude art, business, economics, history, philosophy, political science, psychology, and sociology, in order to focus on how AI is changing the natural sciences and career trajectories in the sciences. Our 6 natural science disciplines include the majority of OpenAlex articles, resulting in 41,298,433 papers, containing 18,392,040 in biology, 4,209,771 in chemistry, and 2,380,666 in geology, 4,755,717 in materials science, 24,315,342 in medicine, 5,138,488 in physics. The selected disciplines cover various dimensions of natural science, representing an broad view of scientific research as a whole.
M2. Divide three stages of AI development We divide the history of AI development into three key eras: the traditional machine learning (ML) era (1980–2014), the deep learning (DL) era (2015–2022), and the generative (GAI) era (2023–present). We consider 1980 as the start of the traditional machine learning era because several landmark researches were published in the 1980s, such as the back-propagating method55,56. We regard that the deep learning era began in 2015, as marked by breakthroughs including ResNet, which enabled the training of ultra-deep neural networks, revolutionizing fields including computer vision and speech recognition57. Finally, we divide the GAI era to begin in 2023 with the publication of ChatGPT, a representative large language model, in December 2022, which saw the advent of large-scale transformer-based models capable of strong generalized performance across a wide range of tasks, sparking new applications in natural language processing and beyond.
Discussion. In this study, we perform a large-scale empirical measurement of the effect of adopting AI in science on both individual scientists and scientific communities. We identify three waves of AI adoptions in science, corresponding with the dominance of machine learning, deep learning, and large language models. Each wave is marked with an accelerated AI adoption rate in research papers and authors. We find that individual scientists are increasingly rewarded with expanded academic impact and accelerated career development for incorporating AI assistance in research across these waves and in all natural science research fields we studied. On average, the use of AI helps individual scientists publish 3.02 times more papers, receive 4.84 times more citations, and become team leaders 1.37 years earlier. This substantial academic benefit may be a driving force behind the accelerated rate of AI adoption. However, we also find unintended consequences from the increased prevalence of AI-augmented research. In all fields, AI-augmented research focuses on a narrower scope of scientific topics and reduces the scientific engagement of follow-on research, leading to more overlapping research works that slows the expansion of knowledge. Further, with a greater concentration of collective attention to the same AI papers, the adoption of AI appears to induce authors to engage in collective hill-climbing1, catalyzing solutions to known problems rather than creating new ones. These findings raise critical questions for science policy. What are the topics most likely left behind from AI-augmented research across fields? Those with less available data include critical scientific questions regarding the origins of natural phenomena, where data are necessarily reduced. Accelerating scientific activity “under the lamp post” of highly visible, data-rich phenomena moves science away from many foundational questions and towards operational ones. By driving attention toward the most popular new developments, AI appears to drive problem solution over generation. These issues become particularly concerning in the face of calls to further increase support for AI-augmented science47,48, coupled with the personal scientific incentives we demonstrate. This could shift collective attention away from new and original questions that lack the data required for AI to demonstrate benefit. It is true that more overlapping attention and a contracted focus may benefit scientific replication and extension, accelerating the emergence of solid and practical solutions to specific questions. Insofar as scientific discovery represents a vast and complex landscape, however, concentrating attention on the same developments may increase the likelihood that science becomes fixed on local maxima of scientific explanation and prediction rather than searching in a more broad, decoupled, and diverse way. While our analysis provides new insight into AI’s impact on science, clear limitations remain. Our identification approach, though validated by experts, misses subtle and unmentioned forms of AI use, and our focus on natural sciences excludes important domains where AI adoption patterns may differ. Moreover, despite consistently suggestive evidence, we cannot fully identify the causal linkage between AI adoption and scientific impact. Nevertheless, our findings demonstrate that currently attributed uses of AI in science primarily augment cognitive tasks through data processing and pattern recognition. Looking forward, these findings illuminate a critical and expansive pathway for AI development in science.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Does AI deployment reduce or exacerbate workplace inequality and income instability? What human oversight must AI research systems have?- Do humans or AI perform better at different research stages?
- How does data availability shape which scientific questions AI systems tackle?
- Where does AI assistance become reliable versus prone to failure in science?
- How should labs measure their own AI systems' impact on research workflows?
- Does AI research acceleration compound into faster field-wide progress over time?
- How much can computational speed and automation substitute for human scientific judgment?
- How does automating research tasks change the pace of AI progress?
- How do technological spillovers between research sectors compound growth rates?
- How much sector-level productivity spillover does real AI research exhibit?
- Why do AI-augmented researchers engage less with one another across topics?
- Does narrowing scientific focus toward data-rich problems create long-term research risks?
- Can human-AI collaboration preserve scientific breadth while improving individual productivity?
- Why do early-career researchers adopt AI tools at higher rates?
- How does rising researcher count relate to declining output per scientist?
- How do co-authorship patterns alone capture scientist expertise without surveying reasoning?