SYNTHESIS NOTE
Topics›Domain Specialization›this note

How much peer review text shows signs of LLM modification?

Researchers analyzed AI conference reviews to estimate what fraction might have been substantially altered by large language models. Understanding this helps clarify how AI tools are entering academic peer review.

Synthesis note · 2026-10-06 · sourced from Domain Specialization

The authors estimate that between 6.5% and 16.9% of text submitted as peer reviews to ICLR 2024, NeurIPS 2023, CoRL 2023 and EMNLP 2023 "could have been substantially modified by LLMs, i.e. beyond spell-checking or minor writing updates." The number is a population share for those four venues, not a count of reviews. The authors report no comparable change in Nature family journal reviews. Where they locate user behavior is in the circumstances: the fraction is higher in reviews that report lower confidence, were submitted close to the deadline, and come from reviewers "less likely to respond to author rebuttals."

The method is what the paper calls distributional GPT quantification. It starts from a stated problem: LLM detectors have "unstable performance," so instead of classifying each review and counting, the authors fit a maximum likelihood estimate of the mixture. Reviews known to be human-written supply one reference distribution, and reviews an LLM generates from the same review instructions and papers supply the other. The vocabulary is adjectives, because a word like "commendable" spikes in recent ICLR reviews and occurs more often in generated text, with technical keywords removed. The authors report that adverbs, verbs and nouns give similar results. They also claim the approach is "more than 10 million times" cheaper computationally than state-of-the-art detectors, with estimation error reduced by factors of 3.4 in-distribution and 4.6 out-of-distribution. Those are the authors' own comparisons.

Set against the nearest notes, this is the peer-review case of the population move in How fast did LLM writing adoption actually spread?, which measures four public-facing domains. Its reason for avoiding per-document judgment matches Can people reliably spot content made by AI?, though this excerpt tests no human judges. The corpus-level compression the authors describe, where generated text narrows "linguistic variation and epistemic diversity," is the aggregate form of the lexical gap in Can human judges detect measurable differences in AI text?. The reviewing setting also meets How much does rhetorical style shift AI review scores? from the other side: there, AI reviewers' scores move with rhetoric, while here human reviewers' text is partly LLM-modified. The excerpt does not connect the two.

The excerpt does not establish that reviewers wrote reviews with ChatGPT from scratch. The authors say so directly: "Our method does not constitute direct evidence that reviewers are using ChatGPT to write reviews from scratch." A reviewer who sketches bullet points and has an LLM flesh them out would also produce a high estimate. The estimate rests on one generator family; the authors report that GPT-3.5 data generalizes to GPT-4. The discussion states the share as "roughly 7-15% of sentences," a different unit from the abstract's share of text. The confidence, deadline and rebuttal links are associations, not tested causes. The implication is narrow: for four venues in 2023 and 2024 and one generator family, a corpus-level share of 6.5% to 16.9% is consistent with substantial LLM modification, and the corpus trend is a firmer claim than any label on a single review.

Inquiring lines that read this note 12

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can AI systems perform peer review as effectively as humans? Do restrictions on reviewer LLM use actually shape peer review behavior? How can we detect and account for LLM involvement in academic writing?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 83 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

between 6.5% and 16.9% of AI conference peer review text could be substantially LLM-modified — a corpus estimate that classifies no single review