Can large language models provide useful feedback on research papers? A large-scale empirical analysis
Expert feedback lays the foundation of rigorous research. However, the rapid growth of scholarly production and intricate knowledge specialization challenge the conventional scientific feedback mechanisms. High-quality peer reviews are increasingly difficult to obtain. Researchers who are more junior or from under-resourced settings have especially hard times getting timely feedback. With the breakthrough of large language models (LLM) such as GPT-4, there is growing interest in using LLMs to generate scientific feedback on research manuscripts. However, the utility of LLM-generated feedback has not been systematically studied. To address this gap, we created an automated pipeline using GPT-4 to provide comments on the full PDFs of scientific papers. We evaluated the quality of GPT-4’s feedback through two large-scale studies. We first quantitatively compared GPT-4’s generated feedback with human peer reviewer feedback in 15 Nature family journals (3,096 papers in total) and the ICLR machine learning conference (1,709 papers). The overlap in the points raised by GPT-4 and by human reviewers (average overlap 30.85% for Nature journals, 39.23% for ICLR) is comparable to the overlap between two human reviewers (average overlap 28.58% for Nature journals, 35.25% for ICLR). The overlap between GPT-4 and human reviewers is larger for the weaker papers (i.e., rejected ICLR papers; average overlap 43.80%). We then conducted a prospective user study with 308 researchers from 110 US institutions in the field of AI and computational biology to understand how researchers perceive feedback generated by our GPT-4 system on their own papers. Overall, more than half (57.4%) of the users found GPT-4 generated feedback helpful/very helpful and 82.4% found it more beneficial than feedback from at least some human reviewers. While our findings show that LLM-generated feedback can help researchers, we also identify several limitations. For example, GPT-4 tends to focus on certain aspects of scientific feedback (e.g., ‘add experiments on more datasets’), and often struggles to provide in-depth critique of method design.
Introduction. In the 1940s, Claude Shannon, while at Bell Laboratories, embarked on developing a mathematical framework of information and communication1. Throughout this pursuit, he was faced with the challenge of naming his novel measure and considered terms such as ‘information’ and ‘uncertainty’. Shannon shared his work with John von Neumann, who quickly recognized the profound links between Shannon’s work and statistical mechanics, and proposed what later anchored modern information theory: ‘Information Entropy’2. Scientific progress often rests on feedback and critique. Effective feedback among peer scientists not only elucidates and promotes the way new discoveries are made, interpreted, and communicated, but also catalyzes the emergence of new scientific paradigms by connecting individual insights, coordinating concurrent lines of thoughts, and stimulating constructive debates and disagreement3. However, the process of providing timely, comprehensive, and insightful feedback on scientific research is often laborious, resource-intensive, and complex4. This complexity is exacerbated by the exponential growth in scholarly publications and the deepening specialization of scientific knowledge5,6. Traditional avenues, such as peer review and conference discussions, exhibit constraints in scalability, expertise accessibility, and promptness. For instance, it has been estimated that peer review – one of the most major channels of scientific feedback – costs over 100M researcher hours and $2.5B US dollars in a single year7. Yet at the same time, it has been increasingly challenging to secure enough qualified reviewers who can provide high-quality feedback given the rapid growth in the number of submissions8–12. For example, the number of submissions to the ICLR machine learning conference increased from 960 in 2018 to 4,966 in 2023. While shortage of high-quality feedback presents a fundamental constraint on the sustainable growth of science overall, it also becomes a source of deepening scientific inequalities. Marginalized researchers, especially those from non-elite institutions or resource-limited regions, often face disproportionate challenges in accessing valuable feedback, perpetuating a cycle of systemic scientific inequality13,14. Given these challenges, there is an urgent need for crafting scalable and efficient feedback mechanisms that can enrich and streamline the scientific feedback process. Adopting such advancements holds the promise of not just elevating the quality and scope of scientific research, given the concerning deceleration in scientific advancements15,16, but also of democratizing its access across the scientific community. Large language models (LLMs)17–19, especially those powered by Transformer-based architectures and pretrained at immense scales, have opened up great potential in various applications20–23. While LLMs have made remarkable strides in various domains, the promises and perils of leveraging LLMs for scientific feedback remain largely unknown. Despite recent attempts that explore the potential uses of such tools in areas such as automating paper screening24, error identification25, and checklist verification26 1, we lack large-scale empirical evidence on whether and how LLMs may be used to facilitate scientific feedback and augment current academic practices. In this work, we present the first large-scale systematic analysis characterizing the potential reliability and credibility of leveraging LLM for generating scientific feedback. Specifically, we developed a GPT-4 based scientific feedback generation pipeline that takes the raw PDF of a paper and produces structured feedback (Fig. 1a). The system is designed to generate constructive feedback across various key aspects, mirroring the review structure of leading interdisciplinary journals27,28 and conferences29–33, including: 1) Significance and novelty, 2) Potential reasons for acceptance, 3) Potential reasons for rejection, and 4) Suggestions for improvement. To characterize the informativeness of GPT-4 generated feedback, we conducted both a retrospective analysis and a prospective user study. In the retrospective analysis, we applied our pipeline on papers that had previously been assessed by human reviewers. We then compared the LLM feedback with the human feedback. We assessed the degree of overlap between key points raised by both sources to gauge the effectiveness and reliability of LLM feedback. Furthermore, we compared the topic distributions of LLM feedback and human feedback. To enable such analysis, we curated two complementary datasets containing full-text of papers, their meta information, and associated peer reviews after 2022 2. The first dataset was sourced from Nature family journals, which are leading scientific journals covering multidisciplinary fields including biomedicine and basic sciences.
Method. Generating Scientific Feedbacks using LLM We prototyped a pipeline to generate scientific feedback using OpenAI’s GPT-419 (Fig. 1a). The system’s input was the academic paper in PDF format, which was then parsed with the machine-learning-based ScienceBeam PDF parser53. Given the token constraint of GPT-4, which allows 8,192 tokens for combined input and output, the initial 6,500 tokens of the extracted title, abstract, figure and table captions, and main text were utilized to construct Retrospective Extraction and Matching of Comments from Scientific Feedback To evaluate the overlap between LLM feedback and human feedback, we developed a two-stage comment matching pipeline (Supp. Fig. 6). In the first stage, we employed an extractive text summarization approach34–37. Each feedback text, either from the LLM or a human, was processed by GPT-4 to extract a list of the points of comments raised in the text (see prompt in Supp. Fig. 13). The output was structured in a JSON (JavaScript Object Notation) format. Within this format, each JSON key assigns an ID to a specific point, while the corresponding value details the content of the point (Supp. Fig. 13). We focused on criticisms in the feedback, as they provide direct feedback to help authors improve their papers54. The second stage focused on semantic text matching38–40. Here, we input both the JSON-formatted feedback from the LLM and the human into GPT-4. The LLM then generated another JSON output where each key identified a pair of matching point IDs and the associated value provided the explanation for the match. Given that our preliminary experiments showed GPT-4’s matching to be lenient, we introduced a similarity rating mechanism. In addition to identifying corresponding pairs of matched comments, GPT-4 was also tasked with self-assessing match similarities on a scale from 5 to 10 (Supp. Fig. 14). We observed that matches graded as “5. Somewhat Related” or “6. Moderately Related” introduced variability that did not always align with human evaluations. Therefore, we only retained matches ranked “7. Strongly Related” or above for subsequent analyses. We validated our retrospective comment matching pipeline using human verification. In the extractive text summarization stage, we randomly selected 639 pieces of scientific feedback, including 150 from the LLM and 489 from human contributors. Two co-authors assessed each feedback and its corresponding list of extracted comments, identifying true positives (correctly extracted comments), false negatives (missed relevant comments), and false positives (incorrectly extracted or split comments). This process resulted in an F1 score of 0.968, with a precision of 0.977 and a recall of 0.960 (Supp. Table 3a), demonstrating the accuracy of the extractive summarization stage. For the semantic text matching stage, we sampled 760 pairs of scientific feedbacks: 332 comparing GPT to Human feedback and 428 comparing Human feedbacks. Each feedback pair was processed to enumerate all potential pairings of their extracted comment lists, resulting in 12,035 comment pairs. Three co-authors independently determined whether the comment pairs matched, without referencing the pipeline’s predictions. Comparing these annotations with pipeline outputs yielded an F1 score of 0.824, a recall of 0.878, and a precision of 0.777 (Supp. Table 3b). To assess inter-annotator agreement, we collected three annotations for 800 randomly selected comment pairs. Given the prevalence of non-matches, we employed stratified sampling, drawing 400 pairs identified as matches by the pipeline and 400 as non-matches. We then calculated pairwise agreement between annotations and the F1 score for each annotation against the majority consensus. The data showed 89.8% pairwise agreement and an F1 score of 88.7%, indicating the reliability of the semantic text matching stage.
Evaluating Specificity of LLM Feedback through Review Shuffling To evaluate the specificity of the feedback generated by the LLM, we compared the overlap between human-authored feedback and shuffled LLM feedback.
Discussion. LLM feedback significantly overlaps with human-generated feedback We began by examining the overlap between LLM feedback and human feedback on Nature family journal data (Supp. Table 1). More than half (57.55%) of the comments raised by GPT-4 were raised by at least one human reviewer (Supp. Fig. 1a). This suggests a considerable overlap between LLM feedback and human feedback, indicating potential accuracy and usefulness of the system. When comparing LLM feedback with comments from each individual reviewer, approximately one third (30.85%) of GPT-4 raised comments overlapped with comments from an individual reviewer (Fig. 2a). The degree of overlap between two human reviewers was similar (28.58%), after controlling for the number of comments (Methods). Results were consistent across other overlapping metrics including Szymkiewicz–Simpson overlap coefficient, Jaccard index, Sørensen–Dice coefficient (Supp. Fig. 2). This indicates that the overlap between LLM feedback and human feedback is comparable to the overlap observed between two human reviewers. We further stratified these overlap results by academic journals (Fig. 2c). While the degree of overlap between LLM feedback and human comments varied across different academic journals within the Nature family — from 15.58% in Nature Communications Materials to 39.16% in Nature — the overlap between LLM feedback and human feedback comments largely mirrored the overlap found between two human reviewers. The robustness of the finding further indicates that scientific feedback generated from LLM is similar to what researchers could get from peer reviewers. In parallel experiments, we investigated the comment overlap between LLM feedback and human feedback on ICLR papers data (Supp. Table 2), and the results were largely similar. A majority (77.18%) of the comments raised by GPT-4 were also raised by at least one human reviewer (Supp. Fig. 1b), indicating considerable overlap between LLM feedback and human feedback. When comparing LLM feedback with comments from each individual reviewer, more than one third (39.23%) of GPT-4 raised comments overlapped with comments from an individual reviewer (Fig. 2b). The overlap between two human reviewers was similar (35.25%), after controlling for the number of comments (Methods, Supp. Fig. 2). We further stratified these overlap results by the decision outcomes of the papers (Fig. 2d). Similar to results over Nature family journals, we found that the overlap between LLM feedback and human feedback comments largely mirrored the overlap found between two human reviewers. In addition, as ICLR dataset includes both accepted and rejected papers, we conducted stratification analysis and found a correlation between worse acceptance decisions and larger overlap in ICLR papers. Specifically, papers accepted with oral presentations (representing the top 5% of accepted papers) have an average overlap of 30.63% between LLM feedback and human feedback comments. The average overlap increases to 32.12% for papers accepted with a spotlight presentation (the top 25% of accepted papers), while rejected papers bear the highest average overlap at 47.09%. A similar trend was observed in the overlap between two human reviewers: 23.54% for papers accepted with oral presentations (top 5% accepted papers), 24.52% for papers accepted with spotlight presentations (top 25% accepted papers), and 43.80% for rejected papers. This suggests that rejected papers may have more apparent issues or flaws that both human reviewers and LLMs can consistently identify.
Limitations. of LLM feedback Study participants also discussed limitations of the current system. The most important limitation is its ability to generate specific and actionable feedback, e.g.
• “Potential Reasons are too vague and not domain specific.” • “GPT cannot provide specific technical areas for improvement, making it potentially difficult to improve the paper.” • “The reviews crucially lacked much in-depth critique of model architecture and design, something actual reviewers would be able to comment on given their likely considerable experience in fields closely related to the focus of the paper.” As such, one future direction to improve the LLM based scientific feedback system is to nudge the system towards generating more concrete and actionable feedback, e.g. through pointing to specific missing work, experiments to add. As one participant nicely summarized:
• “(large languge model generated) reviews were less about the content and more about the testing regime as well as less ML details-focused, but this is okay as it still gave relevant and actionable advice on areas of improvement in terms of paper layout and presenting results.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can AI systems perform peer review as effectively as humans?- Can LLM reviewers catch technical issues that human reviewers miss?
- How did researchers measure whether GPT-4 and human reviewers identified the same issues?
- Does matching reviewer points actually mean the feedback is accurate or correct?
- Why did rejected papers show higher overlap between GPT-4 and human reviewers?
- Could AI feedback work as a substitute for human peer review entirely?
- Does an automated reviewer's output actually match human review accuracy?
- What specific errors did participants report finding in the AI-generated reviews?
- Did adding AI reviews actually change peer review decisions or paper outcomes?
- Do AI reviews depend more on writing style than scientific merit?
- Should AI research papers require dedicated automated review systems instead of existing journals?
- Does review length bias affect acceptance decisions at major conferences?
- How much does reviewer consistency vary across different papers at NeurIPS?
- Can automated review systems catch deep methodological flaws or only surface issues?
- What specific tasks do reviewers use AI for most often?
- Could AI improve peer review rigor and catch human-missed errors?
- Can computational inference scaling catch flaws that human expert reviewers miss?
- Did GPT-4 see the entire paper or only a portion of it?
- Can LLM-generated reference reviews detect machine-written peer review submissions?
- Do reviewer rules about LLM use in peer review actually get followed?
- Do reviewer reward badges risk encouraging lenient or superficial reviews?
- Do stricter AI policies actually change how reviewers score manuscripts?
- Do peer review policies banning LLM use actually change reviewer behavior and decisions?
- What happens when conferences enforce bans or limits on reviewer LLM use?