Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026
LLMs are rapidly reshaping peer review, making it important to understand how reviewers use them in practice and how different LLM-use policies affect review outcomes. We investigate these questions through a randomized experiment and an anonymous post-survey at ICML 2026, a major machine learning conference involving over 24,000 papers and 17,000 reviewers. Reviewers were assigned to either a conservative policy prohibiting all LLM use or a permissive policy allowing limited assistance, with randomization among a subset of main-track papers and reviewers. Policy assignment had near-zero effects on final paper decisions, paper scores, and reviewer confidence, although reviews under the permissive policy were 5.5-7% longer. Post-survey responses (N=1,486) revealed diverse attitudes toward LLMs and substantial noncompliance: 22.5% of conservative-policy reviewers reported using an LLM despite the prohibition, and 36.5% of permissive-policy reviewers reported at least one explicitly disallowed use. We discuss implications for future peer-review policy and tool design.
Introduction. Rapidly growing submission volumes are increasing the burden on peer-review systems and amplifying longstanding concerns about “reviewer fatigue” [67]. At CHI and ICML, two leading conferences in human-computer interaction (HCI) and machine learning (ML), the increase in submissions between 2023 and 2026 was 3,182 to 6,730 (+112%) and 6,538 to 24,661 (+277%), respectively. In response to these pressures and the growing capabilities of large language models (LLMs), it has become common for researchers to incorporate the use of LLMs into their peer review practices [50, 52–54]. Computer Science (CS) conferences, in turn, have begun to publish explicit peer-review LLM policies and explore the use of LLM assistance in the review workflow [1, 2, 7, 22, 33–36, 39, 59, 83].1 (See Section A for an overview of these policies.)
Discussion / Conclusion. 6.1 Implications for the Design of Policies and Tools for Peer Review In this section, we reflect on our findings and discuss promising directions for future peer-review policies and tools. 6.1.1 Design Realistic Policies and Cultivate the Right Environment. Traditionally, reviewing has been a valued, voluntary activity. In a survey of 307 CHI reviewers, Nobarany et al. [61] found that “encouraging high-quality research, giving back to the research community, and finding out about new research” were reviewers’ primary motivations for reviewing. Our post-survey responses echoed these themes, describing science as a “community process” and emphasizing that “the fundamental purpose of peer review [...] is to provide expert, accountable, and independent judgment.” These findings suggest that, with the right culture and environment in place, reviewers would be motivated to conduct high-quality reviews. However, our findings suggest that these conditions are not in place, as evidenced by the high rates of noncompliance in both the Pangram analysis (Section 4.2.3) and post-survey results (Section 5.2.1).
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
What safeguards enable trustworthy AI-assisted scientific peer review at scale?- How much has peer review workload grown at major conferences?
- Can automated AI systems assess novelty as well as human reviewers?
- What counts as a final decision versus an executed revision in research?
- How can automated review scale with the flood of AI-generated papers?
- What accountability structures should replace detection when AI automation increases in peer review?
- At what collaboration level should AI reviewers make final acceptance decisions?
- What discovery accuracy would satisfy the false-alert workload reviewers can tolerate?
- What collaboration model between humans and AI best serves peer review?
- Why should AI research prompts be subject to peer review before use?
- Can statistical filtering plus narrative generation fool academic peer review?
- What makes proof writing and paper writing harder to verify than proof grading?
- What role do multi-dimensional quality frameworks play in assessing arguments versus single-metric approaches?
- Can verification and accountability sustain meaningful human work at scale?
- Why does verification of AI work consistently lag behind AI generation?