Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
Abstract The peer review process in major artificial intelligence (AI) conferences faces unprecedented challenges with the surge of paper submissions (exceeding 10,000 submissions per venue), accompanied by growing concerns over review quality and reviewer responsibility. This position paper argues for the need to transform the traditional one-way review system into a bi-directional feedback loop where authors evaluate review quality and reviewers earn formal accreditation, creating an accountability framework that promotes a sustainable, high-quality peer review system. The current review system can be viewed as an interaction between three parties: the authors, reviewers, and system (i.e. conference), where we posit that all three parties share responsibility for the current problems. However, issues with authors can only be addressed through policy enforcement and detection tools, and ethical concerns can only be corrected through selfreflection. As such, this paper focuses on reforming reviewer accountability with systematic rewards through two key mechanisms: (1) a twostage bi-directional review system that allows authors to evaluate reviews while minimizing retaliatory behavior, (2) a systematic reviewer reward system that incentivizes quality reviewing. We ask for the community’s strong interest in these problems and the reforms that are needed to enhance the peer review process.
Introduction. Artificial intelligence (AI) conferences (e.g. ICML, ICLR, NeurIPS, CVPR, etc) play a crucial role as a premier venue for cutting-edge research dissemination, fostering intellectual discourse and collaboration among researchers. Unlike traditional scientific areas where journals are the primary publication outlets (Kim, 2019), the fast-paced nature of AI research has elevated these conferences to the status of first-tier venues, with their double-blind peer-reviewed papers carrying impact comparable to many prestigious journals (Freyne et al., 2010). However, over recent years, the number of paper submissions to major AI conferences has skyrocketed (Figure 1), overwhelming the traditional peer review system (Tran et al., 2020). With such rapid growth, the authors believe that the current AI conference system is becoming increasingly unsustainable where the hardearned conference’s established reputation is breaking down. To put this more directly, this translates to declined review quality.1 This issue, however, cannot be attributed solely to the reviewers - the responsibility and consequent adverse effects are shared among all three main stakeholders: Authors, Reviewers, and the System.
Definition 1.1 (Authors). Authors are individuals who submit manuscripts to AI conferences for peer review and potential publication.
Definition 1.2 (Reviewers). Reviewers are qualified experts, selected from the author community or invited based on expertise, who evaluate submitted manuscripts.
Definition 1.3 (System). System includes the venue (e.g. ICML) and platforms (i.e. OpenReview 2) that supervise the overall peer review process and AI conferences.
Throughout this position paper, we outline why the responsibility of declined review quality is shared by all three main stakeholders and categorize the challenges as either addressable or non-addressable. We then propose our minimal yet actionable approach that addresses the peer review challenges, positing that just as authors are incentivized to produce high-quality papers, reviewers should be similarly motivated through systematic rewards for quality reviews. Specifically, we advocate for 1. Author Feedback: A two-stage double-blind peer review system where authors evaluate review quality through objective criteria, with safeguards protecting reviewers from potential retaliation when providing critical reviews. 2. Reviewer Rewards: A system-backed reward framework that provides incentives for thorough reviews, where reviewers can accumulate and showcase their review contributions as verifiable academic credentials, creating long-term professional value.
Our paper is structured as follows. Section 2 discusses the root causes of the current problems in AI conference peer reviews. Section 3 suggests our minimal, yet actionable changes that can be made to the current peer review process to restructure the power balance between reviewers and authors. Then, Section 4 provides concrete reward systems for reviewers which need the support of the Systems. Consequently, Section 5 elaborates on the steps needed to make these changes happen and discusses practical challenges of our proposals. Section 6 discusses alternative views from our position. Section 7 includes related works to ours. Lastly, Section 8 concludes our position with future directions to improve the peer review process in AI conferences.
Related work. Our position paper discusses methods to improve peer review quality in AI conferences. In this section, we present related works in four key subsections: studies on review quality by AI conferences, LLMs in peer reviews, reviewer rewards, and other approaches to enhance review quality.
7.1. Studies on Review Quality in AI Conferences NeurIPS has performed several comprehensive studies to analyze their peer review process. In both 2014 and 2021, they performed a review consistency experiment, where 10% of papers submitted to their venue were randomly selected and had them reviewed by two independent groups of reviewers (Cortes & Lawrence, 2021; Beygelzimer et al., Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards 2023). Results showed that 16-23% of papers could have been either accepted or rejected based on the reviewing group, indicating a high randomness in peer review quality. Another directly related work to ours is the study performed in 2022, where NeurIPS asked authors, reviewers, and metareviewers to evaluate the quality of peer reviews after the whole submission process (Goldberg et al., 2025). Among the many notable results, we highlight two key insights that directly relate to our proposals: author-outcome bias and elongated review bias. The author-outcome bias indicates a higher likelihood of authors rating the review to be helpful when the reviews recommended acceptance compared to rejection. This finding motivates the need for our twostage review release mechanism, where authors first evaluate reviews based on their positive aspects and demonstrated understanding of the paper, before seeing the final ratings and critiques. This sequential release of reviews could prevent retaliatory scoring by authors and still obtain feedbacks on the review quality. Similarly concerning, the elongated review bias notes that longer reviews are rated higher than shorter reviews even when the two reviews contain the same information. This observation highlights the need to limit the length of the reviews that are released in the first stage of our two-stage review release method.
7.2. LLMs in Peer Reviews While our position paper addresses the need for an LLM review simply to deter human reviewers from solely relying on LLMs for peer review and to provide authors as a soft reference point for flagging LLM-generated reviews, we believe that the topic of using LLMs as a general peer reviewer needs to be mentioned. Given the ongoing debate in the community, we present a balanced discussion of the potential benefits and limitations of using LLMs in the peer review process.
Method. 3. Redesigning Peer Review: A Two-way Method To address the power imbalance between reviewers and authors, we propose a bi-directional, two-stage review method. Our proposed approach represents a minimal modification to the current double-blind peer review system widely adopted by AI conference venues, making it a low-risk intervention. Importantly, our system works within existing review timelines, ensuring practical implementation while addressing reviewer-author power imbalance issues. As such, we provide a brief description of the current double-blind peer review process and illustrate the changes that we propose.
Current Review System. The paper submission and decision process in major AI conferences (e.g. ICML, NeurIPS) follows the steps illustrated in Figure 2. The process proceeds as follows: (1) Authors submit their papers to the the System (e.g. OpenReview) and papers are matched with reviewers based on the System’s policy (e.g. paper bidding, research expertise); (2) Reviewers evaluate their assigned papers and submit their reviews; (3) All reviews and ratings are simultaneously released to authors; (4) Authors and reviewers engage in discussion to address potential misunderstandings, during which reviewers may adjust their scores based on clarifications; and (5) Meta-reviewers (i.e. area chairs) make final decisions based on both the manuscript Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards and its reviews. This is a high-level overview where details may vary across conferences (Kuznetsov et al., 2024).
Addressing LLM Reviews. To first address LLM reviews, we propose to incorporate LLM-generated peer reviews alongside human reviews during the second stage when reviewers evaluate their assigned papers. This could potentially discourage reviewers from solely relying on LLMs for peer review, as similar AI-generated content would be visible to authors. Importantly, this LLM review would be exclusively visible to authors, who would be explicitly informed of their origin and are not required to engage in discussion with these reviews. While reviewers would be notified about the inclusion of LLM-generated reviews in the process, details about the type of LLM service (e.g. ChatGPT, Claude) and prompts would remain confidential. The inclusion of LLM-generated reviews serves two main purposes: (1) LLM reviews act as a psychological deterrent against the few irresponsible reviewers who might otherwise rely entirely on LLMs for evaluations, as they know the System is already incorporating such automated reviews, and (2) it provides authors with a soft reference point to identify and flag potential LLM-generated reviews, detailed in our second proposal.
Two-Stage Release of Reviews. We propose a sequential release of review content rather than the conventional simultaneous disclosure of all reviews and ratings. Specifically, we divide review content into two distinct sections: section one includes neutral to positive elements, including paper summary, strengths, and clarifying questions, while section two contains more critical parts such as weaknesses and overall ratings. Between these releases, authors evaluate reviews from section one based on two key criteria: (1) the reviewer’s comprehension of the paper and (2) the constructiveness of questions in demonstrating careful reading. During this feedback phase, authors can also flag potential LLM-generated reviews5 by comparing them with the provided LLM reviews. This two-stage disclosure prevents retaliatory scoring while providing the minimal safeguards necessary for a fair review. Once authors complete their feedback, section two is promptly disclosed, and the authors are not allowed to modify their evaluations. Subsequently, the review process follows the conventional workflow.
The author feedback scores on reviews function in two ways. In the long term, these scores can be used for the reviewer reward system, detailed in Section 4. In the short term, the feedback scores are used in the final meta-review stage as additional context in assessing the reviews made on the paper.
Discussion. 5. Discussions 5.1. Need for Gradual Implementations While we believe that our proposal is a minimal and feasible change that is needed to make AI conferences more sustainable, not everyone may feel this way. In this sense, we suggest the following two steps that are needed for the gradual adoption of our proposals. First, a large-scale survey from the authors and reviewers regarding the current peer review system in AI conferences is necessary. The survey should be designed to reflect the different perspectives and responsibilities of each role in the review process, including peer reviewers and meta-reviewers. Understanding the distinct challenges faced by reviewers at different levels will provide a comprehensive understanding of how the peer review system can be improved. Second, we suggest implementing our proposed modification to the review system as a pilot program in smaller-scale tracks or workshops before considering broader adoption. Implementing them in a smaller-scale peer review process could provide practical insights into the feasibility of our proposals, while minimizing disruption to the existing review process.
5.2. Practical Challenges The most practical challenge in implementing our proposed methods is to elicit existing Systems (e.g. venues, OpenReview) to make these changes happen. In particular, the reviewer feedback and reviewer reward systems cannot be initiated without cooperation from multiple venues and Open- Review developers. Making changes to the existing peer review system, which appears to work well on the surface, would require immense effort from multiple parties.
In addition, the question of who would be responsible for implementing these changes’ financial costs needs to be answered. Multiple venues could create a collective fund for implementing these changes, but even very big venues often face budgetary constraints. For instance, ICML 2024 initially planned to provide in-kind compensation to the top 10% of reviewers but had to reduce this to a smaller pool of top reviewers due to financial constraints7. These practical concerns highlight the challenges of implementing the proposed changes to the current peer review process.
5.3. Potential Gaming by Reviewers We recognize that while digital badges leverage positive gamification elements (e.g. participation, motivation), they could potentially incentivize reviewers to optimize for rewards through overly liberal reviewing rather than maintaining review standards. To address these concerns, we should use a carefully designed evaluation metric that rewards quality over leniency. Badges should be provided to reviewers who demonstrate thoroughness, identification of crucial issues, and constructive feedback, but not simply through positivity of the assessments. Ultimately, the success of our proposed badge system depends on aligning reviewer incentives with the academic community’s goal of advancing knowledge through constructive peer evaluation.
5.4. In the Era of LLMs As the era of LLMs came earlier than what most people expected, it seems that most AI conferences are now hurrying to adapt to this sudden change while struggling to maintain their traditional review structures. One of the notable changes we have observed in most AI conference websites is the growing guidelines regarding LLM usage in both pa- 7https://medium.com/@icml2024pc/ reviewing-at-icml-2024-a7aa81169d8c Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards per writing and peer reviewing. For instance, ICLR 2025 has explicitly asked authors to submit their usage of LLMs in the paper writing process. These measures indicate that conferences are only beginning to grapple with the impact of LLMs in academic workflows.
Conclusion. At this moment of heightened interest in AI from both academia and the general public, AI conferences stand at the forefront of AI development, where their significance and impact are beyond measure. Their roles of improving, validating, and disseminating papers are crucial as they serve as a vital channel for conveying important information and technology to society. Throughout this position paper, we have consistently argued that their roles as prestigious and reliable AI conferences have become highly unsustainable from all three main stakeholders: authors, reviewers, and the system itself. Given the practical constraints in regulating authors, we have proposed improvements at the reviewer and system levels. Specifically, we proposed a bi-directional review system where authors are protected with minimal safeguards from unprofessional reviewers, and a systematic reviewer reward system that incentivizes reviewers for their academic service. We have also objectively discussed the potential pros and cons of our proposals. We conclude this position paper with the hope that this work can bring these issues into broader public discourse and inspire future researchers to work on these challenges.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can AI systems perform peer review as effectively as humans?- Can multi-stage AI review pipelines catch scientific flaws better than simple language models?
- Did adding AI reviews actually change peer review decisions or paper outcomes?
- Should rhetorical polish in AI reviews be separated from actual technical accuracy?
- How much of ICLR 2026 peer review was already conducted by AI?
- Do AI reviews depend more on writing style than scientific merit?
- Should AI research papers require dedicated automated review systems instead of existing journals?
- Can workshop acceptance rates reliably measure AI research quality compared to main conferences?
- Does review length bias affect acceptance decisions at major conferences?
- How much does reviewer consistency vary across different papers at NeurIPS?
- Could AI improve peer review rigor and catch human-missed errors?
- Does matching reviewer points actually mean the feedback is accurate or correct?
- Could AI feedback work as a substitute for human peer review entirely?
- Can computational inference scaling catch flaws that human expert reviewers miss?
- What effects do preprint servers have on scientific consensus formation?
- Can institutional statements alone correct misconceptions from unreviewed papers?
- Do AI-generated research reviews score papers higher than human reviewers do?
- How often do researchers suspect peer reviews are written by AI?
- Does AI content in reviews correlate with differences in paper quality control?
- Can human reviewers reliably detect AI-written peer review text by sight?
- What prevents venues from implementing two-way feedback systems at scale?
- Do reviewer reward badges risk encouraging lenient or superficial reviews?
- Do stricter AI policies actually change how reviewers score manuscripts?
- Do conference policies banning LLM use actually reduce AI involvement in reviews?