Do LLM reviewers actually favor LLM-written papers?
Does the apparent bias of LLM-assisted peer reviewers toward LLM-generated papers reflect genuine preferential treatment or an artifact of quality distribution? The answer shapes how we interpret reviewer behavior.
Across more than 125,000 paper-review pairs from ICLR, NeurIPS and ICML, the authors find that LLM-assisted reviews look especially kind to LLM-assisted papers, but that this does not survive a control for paper quality. LLM-assisted reviews are "simply more lenient toward lower quality papers in general," and LLM-assisted papers are over-represented among weaker submissions, which "creates a spurious interaction effect rather than genuine preferential treatment of LLM-generated content." The threshold-varying analysis fits this reading: for the top quality buckets the interaction coefficient is approximately zero (β3 ≈ 0), and the positive interaction in the full corpus comes from the lower buckets.
The mechanism the authors give is rating compression. In their raw comparison (Table I), a human paper's score rises 0.25 points under an LLM-assisted reviewer, against 0.63 points for an LLM-assisted paper, and LLM-assisted reviews "concentrate toward assigning intermediate ratings." Their estimand is β3, the extra benefit of an LLM reviewer for LLM-assisted papers over human-written ones. They check whether interference between reviews could manufacture that term and judge it unlikely: most reviews (75–81%) stay unchanged after the rebuttal, and co-reviewer influence pushes toward consensus. The compression is starkest in fully LLM-generated reviews, which assign scores "almost exclusively in the 6–7 range (on a 1–10 scale) regardless of paper quality." In the fully synthetic condition the kindness differential for LLM-assisted papers is about 0.625 points, against 0.096 points in the observational data, a reduction of roughly 85 percent. Human reviewers using LLMs "substantially reduce this leniency."
This sits against the controlled work in Can LLM judges be fooled by fake credentials and formatting?, where judge biases are shown under prompt manipulation. Here the self-favoring pattern that motivates the paper is tested in observed reviews, and it dissolves once quality is held fixed. The Discussion makes the same point from another angle: "these patterns largely disappear when restricting analysis to accepted papers." The result also qualifies How much does rhetorical style shift AI review scores?. That work shows framing moves LLM scores with content held fixed, a mechanism this excerpt does not test, so presentation sensitivity and quality-driven leniency could both be operating. On policy, the ICML trial in Does banning LLM use in peer review change review outcomes? measured what rules do to scores. The authors' concern is different: a declaration scheme that pairs reviewers and authors by declared LLM use "could inadvertently amplify" any LLM-LLM bias.
The excerpt does not establish why reviewers and metareviewers use LLMs differently. The Limitations section offers rival explanations, such as metareviewers using LLMs for accept decisions but not for rejections, or LLM use concentrating among more lenient reviewers, and it calls for "full-scale user studies" to test them. LLM use is also inferred from a statistical estimate of the LLM-generated fraction, with a 15 percent threshold, so misclassification would blur the comparison. The metareview finding (LLM-assisted metareviews more likely to accept given equal scores) is associational, and the data cover only conferences that publish numerical scores. The implication is narrower than a verdict on LLM bias. Raw score gaps between LLM-assisted and human papers should not be read as favoritism without a quality-conditioned comparison. The finding this excerpt supports most directly is rating compression in fully generated reviews, which the authors also tie to earlier work by Zhu et al.
Inquiring lines that read this note 34
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can AI systems perform peer review as effectively as humans?- Does review length bias affect acceptance decisions at major conferences?
- How much does reviewer consistency vary across different papers at NeurIPS?
- Can LLM reviewers catch technical issues that human reviewers miss?
- Why did rejected papers show higher overlap between GPT-4 and human reviewers?
- Does rhetorical presentation bias reviewers against substantive scientific contributions?
- Does rhetorical quality in reviews influence paper acceptance scores more than content?
- Why do authors submit manuscripts to venues beyond their reach?
- What role do conference organizers play in accepting problematic articles?
- Can LLM-generated reference reviews detect machine-written peer review submissions?
- What quality differences exist between flagged and unflagged peer reviews?
- Do LLM adopters actually cite more diverse and younger research?
- How does LLM-modified writing narrow linguistic diversity in peer review?
- Why do computer science papers show more LLM modification than other fields?
- Can researchers detect individual papers modified by LLMs reliably?
- Did LLM use in scientific writing plateau after the initial surge?
- How much do LLM reviewers shift scores based on rhetorical framing alone?
- How does framing critical topics shape LLM review scores?
- Why do readers rate LLM-edited text more favorably?
- How much does rhetorical framing shift LLM reviewer scores independent of content?
- Can reviewer-author matching by LLM use amplify biases in acceptance decisions?
- What mechanisms drive rating compression in fully LLM-generated peer reviews?
- Do reviewer reward badges risk encouraging lenient or superficial reviews?
- Do peer review policies banning LLM use actually change reviewer behavior and decisions?
- What happens when conferences enforce bans or limits on reviewer LLM use?
- Can rules against undisclosed LLM use change reviewer behavior without enforcement?
- Can watermark-based detection measure true prevalence of LLM use in peer review?
- Do reviewer rules about LLM use in peer review actually get followed?
- Why do peer review policies often fail to change actual review scores?
- Do peer reviewers actually follow policies that ban or limit their LLM use?
- How effective are journal policies restricting LLM use in peer review?
- Can peer review policies actually prevent LLM use when compliance is hard to monitor?
- Do metareviewers and regular reviewers use LLMs differently in peer review?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does banning LLM use in peer review change review outcomes?
Can policies restricting or allowing AI tools shift how reviewers score papers and make decisions? This matters because review quality and fairness depend on consistent standards.
the policy trial found near-zero score effects; this excerpt tests the two-sided LLM interaction and finds it mostly quality confounding.
-
How much does rhetorical style shift AI review scores?
When manuscripts are rewritten to improve rhetoric while keeping scientific content identical, do LLM reviewers change their scores? Understanding this matters for ensuring AI-assisted peer review evaluates substance, not polish.
qualifies: framing moves LLM scores at fixed content, a mechanism this excerpt does not test, so both may operate.
-
Can LLM judges be fooled by fake credentials and formatting?
Explores whether language models evaluating text fall for authority signals and visual presentation unrelated to actual content quality, and whether these weaknesses can be exploited without deep model knowledge.
contrasts: judge biases are shown under controlled prompt attacks; in observed reviews the apparent self-favoring pattern dissolves once quality is controlled.
-
Do LLM reviewers favor papers written by other LLMs?
When LLMs serve as peer reviewers, do they systematically score papers differently based on whether they were written by humans or LLMs? Understanding this matters for fair and trustworthy publication systems.
contradicts: simulated LLM reviewers still favor LLM-written papers, traced to style and criticism aversion, not quality confounding
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
- Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026
- LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on ArXiv
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- A Retrospective on the ICLR 2026 Review Process
- Stop Automating Peer Review Without Rigorous Evaluation
Original note title
the apparent interaction favoring LLM papers in peer review is spurious — LLM-assisted reviews are simply lenient toward lower quality papers