Mapping the Increasing Use of LLMs in Scientific Papers

Paper · arXiv 2404.01268 · Published April 1, 2024
Domain Specialization in LLMs

Scientific publishing lays the foundation of science by disseminating research findings, fostering collaboration, encouraging reproducibility, and ensuring that scientific knowledge is accessible, verifiable, and built upon over time. Recently, there has been immense speculation about how many people are using large language models (LLMs) like ChatGPT in their academic writing, and to what extent this tool might have an effect on global scientific practices. However, we lack a precise measure of the proportion of academic writing substantially modified or produced by LLMs. To address this gap, we conduct the first systematic, large-scale analysis across 950,965 papers published between January 2020 and February 2024 on the arXiv, bioRxiv, and Nature portfolio journals, using a population-level statistical framework to measure the prevalence of LLM-modified content over time. Our statistical estimation operates on the corpus level and is more robust than inference on individual instances. Our findings reveal a steady increase in LLM usage, with the largest and fastest growth observed in Computer Science papers (up to 17.5%). In comparison, Mathematics papers and the Nature portfolio showed the least LLM modification (up to 6.3%). Moreover, at an aggregate level, our analysis reveals that higher levels of LLM-modification are associated with papers whose first authors post preprints more frequently, papers in more crowded research areas, and papers of shorter lengths. Our findings suggests that LLMs are being broadly used in scientific writings.

Introduction. Since the release of ChatGPT in late 2022, anecdotal examples of both published papers (Okunyt ̇e, 2023; Deguerin, 2024) and peer reviews (Oransky & Marcus, 2024) which appear to be ChatGPT-generated have inspired humor and concern.1 While certain tells, such as “regenerate response” (Conroy, 2023b;a) and “as an AI language model” (Vincent, 2023), found in published papers indicate modified content, less obvious cases are nearly impossible to detect at the individual level (Else, 2023; Gao et al., 2022). Liang et al. (2024) present a method for detecting the percentage of LLM-modified text in a corpus beyond such obvious cases. Applied to scientific publishing, the importance of this at-scale approach is two-fold: first, rather than looking at LLM-use as a type of rule-breaking on an individual level, we can begin to uncover structural circumstances which might motivate its use. Second, by examining LLM-use in academic publishing at-scale, we can capture epistemic and linguistic shifts, miniscule at the individual level, which become apparent with a birdseye view.

Measuring the extent of LLM-use on scientific publishing has urgent applications. Concerns about accuracy, plagiarism, anonymity, and ownership have prompted some prominent scientific institutions to take a stance on the use of LLM-modified content in academic publications. The International Conference on Machine Learning (ICML) 2023, a major machine learning conference, has prohibited the inclusion of text generated by LLMs like ChatGPT in submitted papers, unless the generated text is used as part of the paper’s experimental analysis (ICML, 2023). Similarly, the journal Science has announced an update to their editorial policies, specifying that text, figures, images, or graphics generated by ChatGPT or any other LLM tools cannot be used in published works (Thorp, 2023). Taking steps to measure the extent of LLM-use can offer a first-step in identifying risks to the scientific publishing ecosystem. Furthermore, exploring the circumstances in which LLMuse is high can offer publishers and academic institutions useful insight into author behavior. Sites of high LLM-use can act as indicators for structural challenges faced by scholars. These range from pressures to “publish or perish” which encourage rapid production of papers to concerns about linguistic discrimination that might lead authors to use LLMs as prose editors.

We conduct the first systematic, large-scale analysis to quantify the prevalence of LLMmodified content across multiple academic platforms, extending a recently proposed, stateof-the-art distributional GPT quantification framework (Liang et al., 2024) for estimating the fraction of AI-modified content in a corpus. Throughout this paper, we use the term “LLMmodified” to refer to text content substantially updated by ChatGPT beyond basic spelling and grammatical edits. Modifications we capture in our analysis could include, for example, summaries of existing writing or the generation of prose based on outlines.

Related work. GPT Detectors Various methods have been proposed for detecting LLM-modified text, including zero-shot approaches that rely on statistical signatures characteristic of machinegenerated content (Lavergne et al., 2008; Badaskar et al., 2008; Beresneva, 2016; Solaiman et al., 2019; Mitchell et al., 2023a; Yang et al., 2023a; Bao et al., 2023; Tulchinskii et al., 2023) and training-based methods that finetune language models for binary classification of human vs. LLM-modified text (Bhagat & Hovy, 2013; Zellers et al., 2019; Bakhtin et al., 2019; Uchendu et al., 2020; Chen et al., 2023; Yu et al., 2023; Li et al., 2023; Liu et al., 2022; Bhattacharjee et al., 2023; Hu et al., 2023a). However, these approaches face challenges such as the need for access to LLM internals, overfitting to training data and language models, vulnerability to adversarial attacks (Wolff, 2020), and bias against non-dominant language varieties (Liang et al., 2023a). The effectiveness and reliability of publicly available LLM-modified text detectors have also been questioned (OpenAI, 2019; Jawahar et al., 2020; Fagni et al., 2021; Ippolito et al., 2019; Mitchell et al., 2023b; Gehrmann et al., 2019; Heikkil ̈a, 2022; Crothers et al., 2022; Solaiman et al., 2019; Kirchner et al., 2023; Kelly, 2023), with the theoretical possibility of accurate instance-level detection being debated (Weber-Wulff et al., 2023; Sadasivan et al., 2023; Chakraborty et al., 2023). In this study, we apply the recently proposed distributional GPT quantification framework (Liang et al., 2024), which estimates the fraction of LLM-modified content in a text corpus at the population level, circumventing the need for classifying individual documents or sentences and improving upon the stability, accuracy, and computational efficiency of existing approaches. A more comprehensive discussion of related work can be found in Appendix G.

Method. A key characteristic of this framework is that it operates on the population level, without the need to perform inference on any individual instance. As validated in the prior paper, the framework is orders of magnitude more computationally efficient and thus scalable, produces more accurate estimates, and generalizes better than its counterparts under significant temporal distribution shifts and other realistic distribution shifts.

We apply this framework to the abstracts and introductions (Figures 1 and 7) of academic papers across multiple academic disciplines,including arXiv, bioRxiv, and 15 journals within the Nature portfolio, such as Nature, Nature Biomedical Engineering, Nature Human Behaviour, and Nature Communications. Our study analyzes a total of 950,965 papers published between January 2020 and February 2024, comprising 773,147 papers from arXiv, 161,280 from bioRxiv, and 16,538 from the Nature portfolio journals. The papers from arXiv cover multiple academic fields, including Computer Science, Electrical Engineering and Systems Science, Mathematics, Physics, and Statistics. These datasets allow us to quantify the prevalence of LLM-modified academic writing over time and across a broad range of academic fields.

We adapt the distributional LLM quantification framework from Liang et al. (2024) to quantify the use of AI-modified academic writing. The framework consists of the following steps:

  1. Problem formulation: Let P and Q be the probability distributions of human-written and LLM-modified documents, respectively. The mixture distribution is given by Dα(X) = (1 −α)P(x) + αQ(x), where α is the fraction of AI-modified documents. The goal is to estimate α based on observed documents {Xi}N i=1 ∼Dα. 2. Parameterization: To make α identifiable, the framework models the distributions of token occurrences in human-written and LLM-modified documents, denoted as PT and QT, respectively, for a chosen list of tokens T = {ti}M i=1. The occurrence probabilities of each token in human-written and LLM-modified documents, pt and qt, are used to parameterize PT and QT:

Liang et al. (2024) demonstrate that the data points {Xi}N i=1 ∼Dα can be constructed either as a document or as a sentence, and both work well. Following their method, we use sentences as the unit of data points for the estimates for the main results. In addition, we extend this framework for our application to academic papers with two key differences:

Generating Realistic LLM-Produced Training Data We use a two-stage approach to generate LLM-produced text, as simply prompting an LLM with paper titles or keywords would result in unrealistic scientific writing samples containing fabricated results, evidence, and ungrounded or hallucinated claims.

Specifically, given a paragraph from a paper known to not include LLM-modification, we first perform abstractive summarization using an LLM to extract key contents in the form of an outline. We then prompt the LLM to generate a full paragraph based the outline (see Appendix for full prompts).

Our two-stage approach can be considered a counterfactual framework for generating LLM text: given a paragraph written entirely by a human, how would the text read if it conveyed almost the same content but was generated by an LLM? This additional abstractive summarization step can be seen as the control for the content. This approach also simulates how scientists may be using LLMs in the writing process, where the scientists first write the outline themselves and then use LLMs to generate the full paragraph based on the outline.

Using the Full Vocabulary for Estimation We use the full vocabulary instead of only adjectives, as our validation shows that adjectives, adverbs, and verbs all perform well in our application (Figure 3). Using the full vocabulary minimizes design biases stemming from vocabulary selection. We also find that using the full vocabulary is more sample-efficient in producing stable estimates, as indicated by their smaller confidence intervals by bootstrap.

Discussion. Our findings show a sharp increase in the estimated fraction of LLM-modified content in academic writing beginning about five months after the release of ChatGPT, with the fastest growth observed in Computer Science papers. This trend may be partially explained by Computer Science researchers’ familiarity with and access to large language models. Additionally, the fast-paced nature of LLM research and the associated pressure to publish quickly may incentivize the use of LLM writing assistance (Foster et al., 2015).

Preprint more likely to rely on AI for writing assistance. These results may be an indicator of the competitive nature of certain research areas and the pressure to publish quickly.

Limitations. If the majority of modification comes from an LLM owned by a private company, there could be risks to the security and independence of scientific practice. We hope our results inspire further studies of widespread LLM-modified text and conversations about how to promote transparent, epistemically diverse, accurate, and independent scientific publishing.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Can AI systems perform peer review as effectively as humans? Do restrictions on reviewer LLM use actually shape peer review behavior? How can we detect and account for LLM involvement in academic writing? What prevents LLMs from applying their reasoning knowledge to improve outputs? How do hallucinated citations emerge in AI scholarly output? How reliably can humans and AI detectors identify machine-generated text? Does disclosing AI authorship change how audiences evaluate the writing? Can readers reliably distinguish AI-written text from human writing?