Pangram Predicts 21% of ICLR Reviews are AI-Generated

Paper · Source
Domain Specialization in LLMs

Source: Pangram Labs · 2025-11-18

Are authors using LLMs to write AI research papers? Are peer reviewers outsourcing the writing of their reviews of these papers to generative AI tools? In order to find out, we analyzed all 19,000 papers and 70,000 reviews from the International Conference on Learning Representations, one of the most important and prestigious AI research publication venues. Thanks to OpenReview and ICLR's public review process, all of the papers and their reviews were made publicly available online, and this open review process enabled this analysis.

In all seriousness, many ICLR authors and reviewers have been noticing some cases of blantant AI-related scientific misconduct, such as an LLM-generated paper with completely hallucinated references, and many authors claiming to receive completely AI-generated reviews.

One author even reported that a reviewer asked 40 AI-generated questions in their peer review!

We wanted to measure the scale of this problem at large: are these examples of bad behavior one-off incidents, or are they indicative of a larger pattern at work? That's why we took Graham up on his offer!

So, we do not perform this study as a means of calling out individual offenders- as LLMs are actually allowed in both the paper submission and the peer review process. We instead wish to draw attention to the amount of AI usage in the papers and peer review, and highlight that fully AI-generated reviews (which indeed, are likely to be Code of Ethics violations) are a much more widespread problem than many realize.

We found 21%, or 15,899 reviews, were fully AI-generated. We found over half of the reviews had some form of AI involvement, either AI editing, assistance, or full AI-generation.

Paper submissions, on the other hand, are still mostly human-written (61% were mostly human-written). However, we did find several hundred fully AI-generated papers, though they seem to be outliers, and 9% of submissions had over 50% AI content. As a caveat, some fully AI-generated papers were already desk rejected and removed from OpenReview before we had a change to perform the analysis.

Contrary to a previous study that showed that LLMs often prefer their own outputs to human writing when used as a judge, we find the opposite: the more AI-generated text present in a submission, the worse the reviews are.

We find the more AI is present in a review, the higher the score is. This is problematic: it means rather than reframing the reviewer's own opinion using AI as the frame (if this were the case, we'd expect the average score to be the same for AI reviews and human reviews), reviewers are actually outsourcing the judgement of the paper to AI as well. Misrepresenting the LLM's opinion as a reviewer's own actual opinion is a clear violation of the Code of Ethics. We know that AI tends to be sycophantic, which means it says things that people want to hear and are pleasing rather than giving an unbiased opinion: a completely undesirable property when applied to peer review! This could explain the positive bias in scores among AI reviews.

Previously a longer review meant that the review was well thought-out and higher quality, but in the era of LLMs, it can often mean the opposite. AI-generated reviews are longer and have a lot of "filler content" in them. According to Shaib et. al., in a research paper called Measuring AI Slop in Text, one property of AI "slop" is that it has low information density-- which means the AI uses a lot of words to say very little in terms of actual content.

We find this to be true in the LLM reviews as well: AI is using a lot of words but not actually giving very high information dense feedback. We argue this is problematic because authors have to waste time parsing a long review and answering vacuous questions that don't actually contain much helpful feedback. It is also worth mentioning that most authors will probably ask a large language model for a review of their submission before they actually submit it. In these cases, the feedback from an LLM review is largely redundant and unhelpful, because the author has already seen the obvious criticisms that an LLM will make.

Pangram's overall false positive rate is 1 in 10,000 on test set documents.

Pangram's false positive rate on held-out scientific papers from ArXiV is 1 in 100,000.

To put these numbers into context, the false positive rate of Pangram is comparable to the false positive rate of DNA testing or a drug test: a true false positive, where a fully AI-generated text is confused with a fully-human text, is non-zero, but exceedingly rare.

The biggest issue with poor quality AI-generated papers is that they simply waste time and resources that are in limited supply. According to our analysis, AI-generated papers are simply not as good as human-written papers, and even more problematically, they can be generated cheaply by dishonest reviewers and paper mills that "spray and pray" (submit a high volume of submissions to a conference in hopes that one of them will get accepted by chance). If AI-generated papers are allowed to flood the peer review system, review quality will continue to decline, and reviewers will be less motivated by having to read "slop" papers instead of real research.

Understanding why AI-generated reviews can be harmful is a bit more nuanced. We agree with ICLR that AI can be used positively in an assistive capacity to help reviewers better articulate their ideas, especially when English is not a reviewer's native language. Additionally, AI can often provide genuinely helpful feedback, and it is often productive for authors to roleplay the peer review process with LLMs, to get the LLMs to critique and poke holes in the research, and catch mistakes and errors that the author may not have caught originally.

However, the question remains: if AI can generate helpful feedback, why should we prohibit fully AI-generated reviews? University of Chicago economist Alex Imas articulates the core issue in a recent tweet: the answer depends on whether we want human judgment involved in scientific peer review.

If we believe current AI models are sufficient to replace human judgment entirely, then conferences should simply automate the entire review process—feed papers through an LLM and assign scores automatically. But if we believe human judgment should remain part of the process, then fully AI-generated content must be sanctioned. Imas identifies two key problems: first, a pooling equilibrium where AI-generated content (being easier to produce) will quickly crowd out human judgment within a few review cycles; and second, a verification problem where determining if an AI review is actually good requires the same effort as reviewing the paper yourself—so if LLMs can generate better reviews than humans, why not automate the entire process?

In my opinion, human judgments are complementary, yet provide orthogonal value to AI reviews. Humans can often come up with out of distribution feedback that may not be immediately obvious. Expert opinions are more useful than LLMs because their opinions are shaped by experience, context, and a perspective that is curated and refined over time. LLMs are powerful, but their reviews often lack taste, judgment, and therefore feel "flat."

Perhaps conferences in the future can put the SOTA LLM review next to the human reviews to ensure that the human reviews are not just restating the "obvious" critiques that can be pointed out by an LLM.

The rise of AI-generated content in academic peer review represents a critical challenge for the scientific community. Our analysis shows that fully AI-generated peer reviews represent a significant proportion of the overall ICLR review population, and the number of AI-generated papers is also rising. Yet, these AI-generated papers are more often slop than genuine research contributions.

We argue that this trend is problematic and harmful for science, and we call on conferences and publishers to embrace AI detection as a solution to deter abuse and preserve scientific integrity.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Can AI systems perform peer review as effectively as humans? What are the real-world consequences of AI citation hallucinations? How do educators verify student capability when AI can produce indistinguishable work? How do hallucinated citations emerge in AI scholarly output? What human oversight must AI research systems have? How can we detect and account for LLM involvement in academic writing? Why do LLM research ideation systems generate novelty but lack diversity?