Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025

Paper · arXiv 2504.09737 · Published April 13, 2025
Domain Specialization in LLMs

Peer review at AI conferences is stressed by rapidly rising submission volumes, leading to deteriorating review quality and increased author dissatisfaction. To address these issues, we developed Review Feedback Agent, a system leveraging multiple large language models (LLMs) to improve review clarity and actionability by providing automated feedback on vague comments, content misunderstandings, and unprofessional remarks to reviewers. Implemented at ICLR 2025 as a large randomized control study, our system provided optional feedback to more than 20,000 randomly selected reviews. To ensure high-quality feedback for reviewers at this scale, we also developed a suite of automated reliability tests powered by LLMs that acted as guardrails to ensure feedback quality, with feedback only being sent to reviewers if it passed all the tests. The results show that 27% of reviewers who received feedback updated their reviews, and over 12,000 feedback suggestions from the agent were incorporated by those reviewers. This suggests that many reviewers found the AI-generated feedback sufficiently helpful to merit updating their reviews. Incorporating AI feedback led to significantly longer reviews (an average increase of 80 words among those who updated after receiving feedback) and more informative reviews, as evaluated by blinded researchers. Moreover, reviewers who were selected to receive AI feedback were also more engaged during paper rebuttals, as seen in longer author-reviewer discussions. This work demonstrates that carefully designed LLM-generated review feedback can enhance peer review quality by making reviews more specific and actionable while increasing engagement between reviewers and authors. The Review Feedback Agent is publicly available at https://github.com/zou-group/review feedback agent.

Introduction. Scientific peer review is a critical step before publication, where domain experts evaluate the research to ensure thoroughness and scientific integrity, prevent false claims, and provide a strong foundation for future work [1, 2]. High-quality reviews are essential for authors to improve their work, address key limitations, and advance scientific progress. However, in a survey of 11,800 researchers worldwide, while 98% view peer review as essential to maintaining the quality and integrity of academic communication, only 55.4% expressed satisfaction with the quality of reviews they receive [3]. This dissatisfaction has grown as obtaining constructive and high-quality peer reviews has become more challenging due to the increase in the number of paper submissions, especially in fast-moving areas like Artificial Intelligence (AI) [4, 5]. For example, the International Conference on Learning Representations (ICLR) experienced year-over-year submission increases of 47% in 2024 and 61% in 2025 [6]. To maintain a rigorous and meaningful peer review process amid this growth, it is crucial to address the growing burden on reviewers and the subsequent deterioration in review quality. Authors at AI conferences increasingly report receiving short, vague reviews with criticisms like ‘not novel’ or ‘not state-of-the-art (SOTA)’ [7]. At the 2023 Association for Computational Linguistics meeting, authors flagged 12.9% of reviews for poor quality, primarily due to these vague, surface-level criticisms [8]. The peer review system is further strained by reviewers being assigned papers outside their expertise [9] and the same papers being reviewed multiple times due to high rejection rates [1]. Additionally, the 2014 NeurIPS Experiment highlighted inconsistencies in the peer review process by showing that approximately 25% of paper acceptance decisions differed between two independent review committees [10]. These issues not only frustrate authors but potentially allow weaker research to be accepted while strong work is rejected, ultimately preventing papers from reaching their full potential due to the decline of meaningful dialogue between reviewers and authors. Large language models (LLMs) [11] have the potential to enhance the quality and usefulness of peer reviews for authors [12]. Recent studies demonstrated that LLMs can serve as effective critics, generating detailed and constructive feedback [13, 14]. Furthermore, LLMs have already shown high utilization in the peer review process. Reviewers are increasingly turning to LLMs to assist in drafting their reviews, with an estimated 10.6% of reviewers at ICLR 2024 using LLMs for this purpose [15, 16]. To explore how LLMs can improve review quality at scale, we introduce Review Feedback Agent, a multi- LLM system designed to enhance the clarity and actionability of reviews by providing feedback to reviewers. Piloted at ICLR 2025 as a large randomized control study, our agent provided feedback to over 20,000 randomly selected reviews (representing half of all ICLR 2025 reviews) over four weeks from October 15 to November 12, 2024. The generated feedback primarily focused on minimizing instances of vague and unjustified comments while also addressing content misinterpretations and unprofessional remarks. Using Claude Sonnet 3.5 as the backbone [11], we created a system of five LLMs that collaborated to generate high-quality feedback. To enhance the system’s reliability against potential errors or failures in instructionfollowing [17, 18], we developed a set of reliability tests to evaluate specific qualities of the generated feedback; the feedback was only posted if it passed all of these tests. Summary of main findings. Of the randomly selected ICLR reviews that received AI feedback, 26.6% of reviewers updated their reviews, altogether incorporating 12,222 suggestions from the feedback agent into the reviews. Blinded ML researchers labeled these revised reviews as more informative and clearer than their initial versions. Reviewers who updated after receiving feedback increased the length of reviews by an average of 80 words. Furthermore, AI feedback led to more engaged discussions during the rebuttal period, as seen through longer author and reviewer responses. We also observed that reviewers who received feedback were more likely to change their scores after the rebuttal period, which was consistent with a more engaged rebuttal process. In this study, we present the first large-scale deployment for using LLMs to assist peer review. By making reviews more actionable and informative, we aim to enhance the peer review experience and promote a more constructive scientific process.

Related work. Due to their extensive capabilities, LLMs are being used across every stage of the peer review process. Reviewers increasingly use LLMs to assist in drafting peer reviews [15, 29, 30]. An estimated 17.5% of authors of Computer Science abstracts on arXiv [31] and 10.6% of reviewers at ICLR 2024 [16] used LLMs for writing assistance. Other studies have shown the potential of LLMs to make the entire review pipeline more efficient across various stages [32, 33, 34, 35] such as writing manuscripts [36], initial quality control [37, 38, 27], and even providing AI-generated instructions for how to write reviews [39]. As peer review workloads continue to increase, LLMs present an opportunity to alleviate some of the burden on human reviewers by providing reviews of submitted manuscripts. In a prospective survey study, 308 researchers from 110 institutions received GPT-4-generated feedback on their papers. Of these, 57.4% found the feedback helpful, and 82.4% felt it was more useful than the feedback provided by at least some human reviewers [12]. Building off of this work, [40] proposed a multi-agent review generation system that improved the specificity and helpfulness of feedback provided compared to GPT-4, reducing the rate of generic comments from 60% to 29%. Furthermore, LLMs offer an efficient and possibly less biased alternative to human evaluations; [41] found that human evaluators of peer reviews were highly susceptible to bias from review length and paper score, as there were high levels of subjectivity among reviewers. These findings suggest that integrating LLMs into the review evaluation process could standardize assessments and reduce inconsistencies. As LLM-based tools continue to evolve, they hold the potential to improve both the speed and quality of manuscript evaluations. Our experiment is the first to demonstrate how LLMs can improve the peer review process on a large scale, highlighting their practical benefits. However, despite these advancements, no prior studies had specifically examined how LLMs could be used to provide feedback on peer reviews in the areas we focused on in our experiment. A study released after our ICLR experiment, however, introduced a benchmark to identify toxicity in peer reviews [42]. The authors identified four categories of toxic comments: using emotive or sarcastic language, vague or overly critical feedback, personal attacks, and excessive negativity. These categories align closely with the ones we chose for our agent to provide feedback on. The authors benchmarked several LLMs for detecting toxicity and tested their ability to revise toxic sentences, finding that human evaluators preferred 80% of these revisions.

Method. In what follows, we first describe the review feedback experiment, including its goals and our technical setup with OpenReview. Next, we outline the architecture of our Review Feedback Agent and explain how the system was designed to meet our goals while ensuring a high level of reliability. In total, the agent automatically provided feedback to over 20, 000 reviews at ICLR 2025.

Our pilot study was conducted in collaboration with ICLR 2025 and OpenReview. As one of the world’s fastest-growing AI conferences, ICLR receives thousands of paper submissions yearly; in 2025, ICLR received 11,603 submissions. Each submission is assigned an average of 4 reviewers, and all reviews are standardized to include the same sections: summary, strengths, weaknesses, and questions. Furthermore, reviewers provide scores on a scale of 1 (low) to 10 (high), rating the paper according to the following categories: soundness, presentation, contribution, rating, and confidence. Goal: Our goal was to enhance review quality and, in particular, reduce low-information content reviews. Toward this goal, we identified three categories of common issues in reviews that we hoped to improve by providing LLM-generated feedback. The common issues are: 1) vague or generic critiques in reviews (the feedback asks the reviewers to be more specific and actionable); 2) questions or confusions that could be addressed by overlooked parts of the paper (the feedback highlights relevant sections); and 3) unprofessional statements in the review (the feedback asks the reviewer to rephrase). For each comment in a review, the Review Feedback Agent determined if it fell into any of these problematic categories and, if so, provided feedback on that specific review comment. Experimental setup: We set up this experiment as a Randomized Control Trial (RCT) to enable us to make causal inferences about how receiving feedback influences the peer review process. Before the beginning of the review period, we randomly split papers into one of three equal groups (see Figure 1A):

  1. No reviews for this paper will receive feedback, 2. Half of the reviews for this paper will be randomly selected to receive feedback, 3. All reviews for this paper will receive feedback.

For reviews randomly assigned to receive feedback, the Review Feedback Agent, wrapped in an API, was automatically triggered when a reviewer first submitted their review on OpenReview. We delayed the feedback generation by one hour after a review was initially submitted to allow reviewers time to make any small edits (e.g., typo corrections). See Figure 1A for an example timeline. The agent posted feedback to reviews through the OpenReview interface by replying to reviews with the feedback wrapped in a comment. See Figure 2 for an example of what feedback looked like on the OpenReview website.

The agent only provided feedback on the initial review, and there was no subsequent interaction between the reviewer and the feedback system after that time point. The feedback is only visible to the reviewer and the ICLR program chairs; it was not shared with other reviewers, authors, or area chairs and was not a factor in the acceptance decisions. Reviewers were informed that the feedback was generated by a LLM and could choose to ignore the feedback or revise their review in response, as the system did not make any direct changes. Finally, we did not access or store any identifiable information about authors or reviewers. This study was reviewed by IRB and deemed low risk. Statistics: Around 50% of reviews were randomly selected to receive feedback. Of the 44,831 reviews submitted on 11,553 unique papers (we excluded desk-rejected submissions), we posted feedback to 18,946 reviews (42.3%) over 4 weeks from October 15 to November 12, 2024 (see Figure 2A). Less than 8% of the selected reviews did not receive feedback for one of two reasons: 2,692 reviews were originally well-written and did not need feedback, while 829 reviews had feedback that failed the reliability tests. Each review took roughly one minute to run through our entire pipeline and cost around 50 cents.

Discussion. Our research demonstrates the significant potential of LLM-based systems to enhance peer review quality at scale. By providing targeted feedback to reviewers at ICLR 2025, we observed meaningful improvements in review specificity, engagement, and actionability. We saw that 27% of reviewers updated their reviews, and an overwhelming majority of those who made updates incorporated at least one piece of feedback into their modifications. Blinded AI researchers found the updated reviews to be consistently more clear and informative. Furthermore, feedback intervention led to increased engagement throughout the review process, with longer reviews, rebuttals, and reviewer responses, suggesting more involved discussions between authors and reviewers. We designed the AI feedback system to enhance reviews while ensuring human reviewers retain complete control. First, the AI-generated feedback was purely optional, and reviewers could decide whether to incorporate it or not; by default, they could opt out by ignoring the feedback. Second, human reviewers had full control over the final review and the scores visible to the authors. To reduce the risk of hallucination, the AI feedback had to pass several rigorous reliability tests before being shared with reviewers. Finally, no personal or identifiable information about reviewers or authors was disclosed to the agent. An IRB review deemed the system to be low risk.

Going forward, there are several directions to further improve the Review Feedback Agent. Our feedback categories focused on three main areas (improving specificity, addressing misunderstandings, and ensuring professionalism). While these categories were derived from reviewer guides and previous studies and encompass the majority of author complaints, they may not capture all aspects of review quality. Expanding to other categories would be helpful. Additionally, it would be interesting to explore the use of reasoning models to generate more nuanced feedback for complex issues in reviews. Finally, the concept of developing reliability tests for LLMs is an evolving field, with new studies emerging after our experiment [43, 44], and we hope to incorporate ideas from these recent works to improve the robustness of our framework. Ultimately, we expect that running this agent at future AI conferences across a diverse range of research topics will improve its robustness and effectiveness. CS conferences have long leveraged machine learning to enhance their peer review processes. One early example is the Toronto Paper Matching algorithm, which was used in NIPS 2010 to match papers with reviewers and has since been deployed by over 50 conferences [45]. However, the impact of many of these earlier applications of machine learning has not been rigorously quantified. To address this gap, we were motivated to conduct this randomized controlled study to rigorously evaluate the effects of review feedback before broader deployment. Our findings show that by striving to make reviews more informative for authors, the Review Feedback Agent has the potential to enhance the overall quality of scientific communication. As LLM capabilities continue to advance, we anticipate even more advanced systems that can provide tailored feedback to reviewers, ultimately benefiting the entire scientific community through improved peer review.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Can AI systems perform peer review as effectively as humans? What human oversight must AI research systems have? How can we detect and account for LLM involvement in academic writing? Do restrictions on reviewer LLM use actually shape peer review behavior?