SYNTHESIS NOTE
Topics›Domain Specialization›this note

Do LLM reviewers actually favor LLM-written papers?

Does the apparent bias of LLM-assisted peer reviewers toward LLM-generated papers reflect genuine preferential treatment or an artifact of quality distribution? The answer shapes how we interpret reviewer behavior.

Synthesis note · 2026-10-06 · sourced from Domain Specialization

Across more than 125,000 paper-review pairs from ICLR, NeurIPS and ICML, the authors find that LLM-assisted reviews look especially kind to LLM-assisted papers, but that this does not survive a control for paper quality. LLM-assisted reviews are "simply more lenient toward lower quality papers in general," and LLM-assisted papers are over-represented among weaker submissions, which "creates a spurious interaction effect rather than genuine preferential treatment of LLM-generated content." The threshold-varying analysis fits this reading: for the top quality buckets the interaction coefficient is approximately zero (β3 ≈ 0), and the positive interaction in the full corpus comes from the lower buckets.

The mechanism the authors give is rating compression. In their raw comparison (Table I), a human paper's score rises 0.25 points under an LLM-assisted reviewer, against 0.63 points for an LLM-assisted paper, and LLM-assisted reviews "concentrate toward assigning intermediate ratings." Their estimand is β3, the extra benefit of an LLM reviewer for LLM-assisted papers over human-written ones. They check whether interference between reviews could manufacture that term and judge it unlikely: most reviews (75–81%) stay unchanged after the rebuttal, and co-reviewer influence pushes toward consensus. The compression is starkest in fully LLM-generated reviews, which assign scores "almost exclusively in the 6–7 range (on a 1–10 scale) regardless of paper quality." In the fully synthetic condition the kindness differential for LLM-assisted papers is about 0.625 points, against 0.096 points in the observational data, a reduction of roughly 85 percent. Human reviewers using LLMs "substantially reduce this leniency."

This sits against the controlled work in Can LLM judges be fooled by fake credentials and formatting?, where judge biases are shown under prompt manipulation. Here the self-favoring pattern that motivates the paper is tested in observed reviews, and it dissolves once quality is held fixed. The Discussion makes the same point from another angle: "these patterns largely disappear when restricting analysis to accepted papers." The result also qualifies How much does rhetorical style shift AI review scores?. That work shows framing moves LLM scores with content held fixed, a mechanism this excerpt does not test, so presentation sensitivity and quality-driven leniency could both be operating. On policy, the ICML trial in Does banning LLM use in peer review change review outcomes? measured what rules do to scores. The authors' concern is different: a declaration scheme that pairs reviewers and authors by declared LLM use "could inadvertently amplify" any LLM-LLM bias.

The excerpt does not establish why reviewers and metareviewers use LLMs differently. The Limitations section offers rival explanations, such as metareviewers using LLMs for accept decisions but not for rejections, or LLM use concentrating among more lenient reviewers, and it calls for "full-scale user studies" to test them. LLM use is also inferred from a statistical estimate of the LLM-generated fraction, with a 15 percent threshold, so misclassification would blur the comparison. The metareview finding (LLM-assisted metareviews more likely to accept given equal scores) is associational, and the data cover only conferences that publish numerical scores. The implication is narrower than a verdict on LLM bias. Raw score gaps between LLM-assisted and human papers should not be read as favoritism without a quality-conditioned comparison. The finding this excerpt supports most directly is rating compression in fully generated reviews, which the authors also tie to earlier work by Zhu et al.

Inquiring lines that read this note 34

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can AI systems perform peer review as effectively as humans? How can we detect and account for LLM involvement in academic writing? Do restrictions on reviewer LLM use actually shape peer review behavior? How can AI systems reliably guide voters without introducing political bias? How can we reduce inherent biases in LLM-based evaluation judges?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 86 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

the apparent interaction favoring LLM papers in peer review is spurious — LLM-assisted reviews are simply lenient toward lower quality papers