Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
Abstract There are increasing indications that LLMs are not only used for producing scientific papers, but also as part of the peer review process. In this work, we provide the first comprehensive analysis of LLM use across the peer review pipeline, with particular attention to interaction effects: not just whether LLM-assisted papers or LLM-assisted reviews are different in isolation, but whether LLM-assisted reviews evaluate LLM-assisted papers differently. In particular, we analyze over 125,000 paper-review pairs from ICLR, NeurIPS, and ICML. We initially observe what appears to be a systematic interaction effect: LLM-assisted reviews seem especially kind to LLM-assisted papers compared to papers with minimal LLM use. However, controlling for paper quality reveals a different story: LLM-assisted reviews are simply more lenient toward lower quality papers in general, and the over-representation of LLM-assisted papers among weaker submissions creates a spurious interaction effect rather than genuine preferential treatment of LLM-generated content. By augmenting our observational findings with reviews that are fully LLM-generated, we find that fully LLM-generated reviews exhibit severe rating compression that fails to discriminate paper quality, while human reviewers using LLMs substantially reduce this leniency. Finally, examining metareviews, we find that LLM-assisted metareviews are more likely to render accept decisions than human metareviews given equivalent reviewer scores, though fully LLM-generated metareviews tend to be harsher. This suggests that meta-reviewers do not merely outsource the decision-making to the LLM. These findings provide important input for developing policies that govern the use of LLMs during peer review, and they more generally indicate how LLMs interact with existing decision-making processes.
Introduction. Since ChatGPT’s release in late 2022, LLMs have become a fixture in research workflows, with 81% of researchers now incorporating them into some aspect of their work (Liao et al. [11]). Recent work has shown evidence of an increasing trend in LLM use by both authors and reviewers at conferences: Liang et al. [9] found that up to 17% of reviews at top machine learning conferences show evidence of LLM modification, while Liang et al. [10] observed that the abstracts of up to 17.5% of recent Computer Science papers showed evidence of LLM use—a significantly higher number than that for papers in other domains. Anecdotally, some researchers have attributed subpar reviews they have received in recent years to LLM use, indicating a general sentiment of distrust in the use of LLMs in peer review. Figure 1a shows the year by year increase in LLM use, in both papers and reviews submitted to the International Conference on Learning Representations (ICLR), by using the technique from Liang et al. [9] to estimate the fraction α of LLM generated text in a document (described in Section 3). Notably, LLM use is substantially more common in writing reviews than in writing papers: in the latest edition of the conference, 3.31% of papers and 26.65% of reviews show substantial LLM modification (detected fraction over αthreshold = 15%). The increase in LLM use in review writing is particularly steep, almost doubling in magnitude in a year. Both LLM-assisted reviews and LLM-assisted papers show interesting differences from those that do not have detected LLM assistance (henceforth referred to as "human"). LLM-assisted papers are frequently given lower scores (Fig. 1b), with ‘3’ being the most commonly assigned quality score, relative to a mode of ‘6’ for human-written papers. Meanwhile LLM assisted reviews concentrate toward assigning intermediate ratings and give fewer harsh ratings (Fig. 1c), suggesting a tendency to hedge rather than provide extreme assessments on either side. An interesting question that naturally arises from LLM use on both sides of the peer review process regards the existence of a systemic interaction effect between LLM/Human papers and LLM/Human Reviews. More directly, does a human reviewer behave differently when presented a fully human paper vs an LLM-assisted paper? Similarly, does an LLM-assisted reviewer behave differently when presented an LLM-assisted paper vs. a fully human paper? Interestingly, ICML has instituted a new policy for their 2026 edition (ICML Policy [2]) requiring both reviewers and authors to declare either conservative or permissive LLM use, with reviewer-author pairing done based on these declarations. If any sort of systematic bias does exist in LLM-LLM interactions within peer review, such a policy could inadvertently amplify it. Moreover, authors could potentially game the system if they understand these interaction effects and their implications for review outcomes. As an initial datapoint, Table I divides (paper, review) pairs into 4 quadrants based on LLM usage—both human-written, human paper with LLM-assisted review, LLM-assisted paper with human review, and both LLM-assisted—and presents the mean and standard deviations of the scores for these pairs. We find that LLM-assisted papers receive lower ratings on average while LLM-assisted reviews give higher ratings. Both human and LLM-assisted papers receive higher scores when reviewed by an LLM-assisted reviewer, but notably the magnitude differs: a 0.25-point jump for human papers versus 0.63 points for LLM-assisted papers. Meanwhile, reviewers show higher confidence for LLM-assisted papers, with human reviewers slightly more confident than LLM-assisted reviewers. These observations suggest potential systematic effects worth exploring, though summary statistics alone cannot verify this due to confounding. It is wholly possible that the papers in one quadrant are fundamentally different papers from the papers in other quadrants, making rating comparisons for these papers moot. This necessitates a causally grounded analysis to eliminate paper-level confounding factors. Motivated by the above observations, our paper uses observational data (paper text, metadata, reviews, scores and metareviews) from top machine learning conferences over the last 3 years, in addition to synthetic data generated by prompting LLMs, to answer the following key research questions:
Are LLM-aided reviews kinder to LLM-aided papers on average?
Does the kindness of LLM-aided reviews to LLM-aided papers vary with paper quality?
How does an LLM-aided review differ from a fully LLM-generated review?
How does LLM-assistance in metareviewing relate to a paper’s final acceptance decision?
Our work differs from prior studies in several key ways. Unlike research that examines LLM use in papers or reviews in isolation, we provide the first comprehensive analysis of interaction effects between LLM use on both sides of the peer-review process.
Related work. The rapid adoption of LLMs in research has motivated the development of detection methods to quantify their use. Liang et al. [9] introduced a statistical approach finding that up to 17% of reviews at top ML conferences showed significant LLM modification, while Liang et al. [10] documented 17.5% of recent Computer Science paper abstracts showed evidence of LLM use. However, they do not study how these distinct forms of LLM use interact with each other. Several studies have examined how LLM-generated reviews differ from human reviews. Zhu et al. [19] identified systematic biases in LLM reviewers, including inflated ratings for weaker papers and reduced sensitivity to paper quality. The susceptibility of LLM reviews to manipulation has been documented by Ye et al. [17] and Keuper [6], who demonstrated successful prompt injection attacks. However, these studies do not explicitly account for LLM-aided reviews. Kocak et al. [7] advocate for LLMs as complementary tools rather than replacements for human judgment. Beyond peer review, several studies have investigated whether LLMs exhibit bias toward their own outputs. Zheng et al. [18] adopt the term ‘self-enhancement bias’ from social cognition literature. Panickssery et al. [14] found evidence that LLMs preferentially rate their own generated text higher, while Laurito et al. [8] extended this to binary choice scenarios, showing LLM-based assistants consistently prefer LLM-presented options. This is an important motivation for our work, and we argue that the setting of peer review serves as a valuable testbed for understanding the implications of two-sided LLM use in practice. OpenReview’s transparency policies provide rich, structured data that allows us to study authentic high-stakes author-reviewer interactions where both parties may have used LLMs.
Method. To address these concerns, we formally model this problem with the causal model in Figure 2. Solid arrows represent hypothesized causal effects; dashed arrows indicate unobserved confounding paths. For each paper-review pair (i, j), we define the treatment Tij ∈{0, 1} as whether review slot j for paper i was assigned to a reviewer that used LLM assistance, Xi as paper covariates including XLLM i ∈{0, 1} (whether paper i contains LLM-assisted text) and Xarea i ∈{0, 1}K (indicator variables for which of K subject areas paper i belongs to), and the outcome Yij as the overall rating (or auxiliary scores like presentation, contribution, or soundness). Let Yij(1) and Yij(0) denote the potential outcomes under treatment and control, respectively. Our key identifying assumption is conditional ignorability:
That is, conditioned on available covariates (in our case, the paper’s area and the paper’s own LLM-assistance status), treatment assignment is independent of potential outcomes. We condition on area because certain subject areas may have both higher rates of LLM-assisted reviewing and systematically different rating distributions, acting as a potential confounder. We are specifically interested in whether the treatment effect of LLM-assisted reviewing differs depending on whether the paper itself was written with LLM assistance. Each paper has multiple review slots, which we treat as interchangeable, and receives some number of unique reviews. As a consequence of OpenReview overwriting original ratings and making them inaccessible after the rebuttal stage, we can only observe the final ratings provided by each review. This raises concerns about violations of the Stable Unit Treatment Value Assumption (SUTVA), which requires that (1) a unit’s potential outcomes depend only on its own treatment assignment, not on the treatment of other units (no interference), and (2) there are no hidden variations in treatment (Rubin [15]). However, several factors suggest these violations are unlikely to create stronger (hence spurious) effects than would exist in pre-rebuttal scores. First, most reviews (75-81%) remain unchanged Kargaran et al. [5], meaning our data predominantly reflects initial, independent assessments. Second, co-reviewer influence tends to push reviews toward consensus Kargaran et al. [5], which would dampen rather than amplify any differential treatment patterns. While access to pre-rebuttal scores would enable stronger causal claims, these considerations suggest post-rebuttal interference is unlikely to generate the interaction effects we document. We now frame the problem as a regression, which allows us to isolate the effects of interest while controlling for potential confounders. We use the conventional way of including interactions in a linear model by modeling the rating Yij for paper i and review j as • β0: The expected rating for a human-written paper receiving a human-written review (the baseline).
• β1: The difference in expected rating for LLM-assisted vs. human-written papers, holding review type constant.
• β2: This is the Conditional Average Treatment Effect (CATE) of receiving an LLM review instead of a human review, for papers written without LLM assistance: CATE(XLLM i = 0) = β2.
• Similarly, the CATE for LLM-assisted papers is given by: CATE(XLLM i = 1) = β2 + β3. This is the additional score an LLM-assisted paper receives if it is reviewed with LLM-assistance instead of entirely by a human.
The interaction coefficient β3 is our primary quantity of interest. It captures the relative leniency of LLM reviews toward LLM-assisted papers, compared to their leniency toward human-written papers. More precisely, it answers the question: how much more (or less) does switching from a human reviewer to a reviewer using LLM assistance benefit an LLM-assisted paper than it benefits a human-written paper?
Discussion. This paper provides empirical insights into how LLM use on both sides of peer review affects ratings and paper outcomes. We find asymmetric patterns: while LLM-aided papers receive lower ratings overall, LLM-aided reviews show stronger positive effects when evaluating LLM-aided papers. Critically, these patterns largely disappear when restricting analysis to accepted papers, indicating effects are driven by lower quality submissions where LLM-aided reviews exhibit general leniency. The degree of rating compression, both in LLM-aided and fully LLM-generated reviews (the latter corroborating Zhu et al. [19]’s findings) is particularly concerning since these reviews ultimately provide less information and introduce noise into the decision process. Beyond review ratings, we find that LLM-aided metareviews are associated with higher acceptance rates despite fully LLM-generated metareviews being overly conservative, and that LLM-aided reviews align less well with final decisions than human reviews, even for borderline papers where decision-making is most consequential. We also make observations about how LLM-aided reviews differ from their fully LLM-generated counterparts, noting the latter’s lack of discriminative power. Periodic reviews of this kind are essential as this landscape continues to evolve. Interesting future directions include comparing our results to those from submissions to similarly reputed journals, which are often less susceptible to some of the issues conferences face: reviewers are not authors competing for the same limited spots, and review timelines are less stressed. We also aim
Conclusion. We return to the research questions introduced in Section 1 with empirically-grounded answers: Are LLM-aided reviews kinder to LLM-aided papers on average and does this vary by paper quality? Our threshold-varying analysis (Figure 4) demonstrates that the interaction effect is strongly modulated by paper quality. For papers in the top quality buckets, we observe minimal differential treatment (β3 ≈0). The positive interaction coefficient observed in the full corpus is driven primarily by papers in the lower quality buckets, where LLM reviews exhibit greater leniency than human reviews, a leniency that benefits LLM-assisted papers more simply because they are overrepresented in this range. As shown in Figure 3, LLM-assisted reviews provide systematically higher ratings to lower quality papers regardless of whether those papers used LLM assistance, but this compression of the rating scale disproportionately affects the (LLM Paper, LLM Review) quadrant due to its compositional skew toward weaker submissions. The within-paper paired analyses with quality buckets (Table VII) confirms this observation. How does an LLM-aided review differ from a fully LLM-generated review? What is the effect of the human in the loop? Our synthetic experiments (Table VIII) reveal substantial differences between unmediated LLM reviews and human-mediated LLM-assisted reviews. Fully LLM-generated reviews exhibit markedly compressed rating distributions, assigning scores almost exclusively in the 6–7 range (on a 1–10 scale) regardless of paper quality (Figure 5). LLM-aided reviews have a broader and more discriminating rating distribution that better tracks paper quality. The kindness differential between LLM and human reviews for LLM-assisted papers is approximately 0.625 points in the fully synthetic condition (ICLR 2025 full corpus) but only 0.096 points in the observational data (Table III), a reduction of roughly 85%. How does LLM use in a metareview relate to a paper’s final decision? The regression analyses in Section 4.4 reveals that given the same set of scores, an LLM-aided metareviewer is more likely to render an accept decision than a human reviewer.
Limitations. However, conferences in other domains (and even journals within CS) do not provide numerical scores with their reviews, which is a major component of our present analysis. Our results on LLM-aided and fully LLM-generated reviews/metareviews together suggest potential patterns in the way people use LLMs in peer review. For instance, metareviewers might selectively deploy LLMs when writing accept decisions, where the summarization task is more straightforward, but write rejection metareviews themselves, where detailed critical reasoning may be required. Alternatively, LLM use could be more popular among metareviewers who are more lenient. Similarly, the gap between LLM-aided reviews and fully LLM-generated reviews could be a result of reviewers asking LLMs to be more critical in their prompts, or they could simply be using LLMs only to expand upon their own salient reviews. We encourage full-scale user studies to test any such hypotheses. Multiple major ML conferences have now instituted LLM policies beyond simple disclosure of use.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Do restrictions on reviewer LLM use actually shape peer review behavior?- Do peer review policies banning LLM use actually change reviewer behavior and decisions?
- What happens when conferences enforce bans or limits on reviewer LLM use?
- Can rules against undisclosed LLM use change reviewer behavior without enforcement?
- Do reviewer rules about LLM use in peer review actually get followed?
- Do peer reviewers actually follow policies that ban or limit their LLM use?
- How effective are journal policies restricting LLM use in peer review?
- Do reviewer reward badges risk encouraging lenient or superficial reviews?
- Can watermark-based detection measure true prevalence of LLM use in peer review?
- Why do peer review policies often fail to change actual review scores?
- Does review length bias affect acceptance decisions at major conferences?
- How much does reviewer consistency vary across different papers at NeurIPS?
- Can LLM reviewers catch technical issues that human reviewers miss?
- Why did rejected papers show higher overlap between GPT-4 and human reviewers?
- Does rhetorical presentation bias reviewers against substantive scientific contributions?
- Does rhetorical quality in reviews influence paper acceptance scores more than content?
- Why do authors submit manuscripts to venues beyond their reach?
- Can LLM-generated reference reviews detect machine-written peer review submissions?
- What quality differences exist between flagged and unflagged peer reviews?
- Do LLM adopters actually cite more diverse and younger research?
- How does LLM-modified writing narrow linguistic diversity in peer review?
- Why do computer science papers show more LLM modification than other fields?
- Can researchers detect individual papers modified by LLMs reliably?
- Did LLM use in scientific writing plateau after the initial surge?
- How much do LLM reviewers shift scores based on rhetorical framing alone?