INQUIRING LINE

An estimated 6.5% to 16.9% of conference peer reviews were substantially changed by AI — does that make them sound alike?

How does LLM-modified writing narrow linguistic diversity in peer review?

This explores whether LLM help with writing peer reviews and papers makes scientific writing sound more alike, and what the corpus can say about how that would happen. The corpus does not directly measure lost linguistic diversity, but it does show how much text is affected and what pressures favor one style.


This explores whether LLM help with writing peer reviews and papers makes scientific writing sound more alike. The corpus doesn't contain a direct measurement of vocabulary or style narrowing in reviews, and it's better to say so than to fake one. What it does have answers two smaller questions: how much review text is touched by LLMs, and what forces push writing toward a single style.

Start with how much text is affected. One analysis of reviews from ICLR, NeurIPS, CoRL and EMNLP estimates that between 6.5% and 16.9% of review text was substantially modified by an LLM How much peer review text shows signs of LLM modification?. The method is worth noticing. It doesn't flag individual reviews. It measures a shift in word patterns across the whole pool of reviews. So the field's best detection tool already treats LLM writing as something that changes the overall shape of the text, not as a feature of single documents. The same study found the most LLM modification among reviewers who were less confident, rushed, or less engaged. Those are the reviews where a reviewer's own voice was least likely to show up anyway.

Why would that text converge? The corpus points to the generation process itself. Token prediction pulls output toward the center of the training data. It doesn't push toward competing positions, so the result is smooth prose that adds claims without adding new perspectives Does LLM generation explore competing claims while producing text?. A related note shows that the same model writes in recognizable registers: a sycophantic chat voice and a 'falsely objective' published-prose voice, each inherited from its training data Why do LLMs produce such different writing in chat versus posts?. A peer review is exactly the kind of text that calls up the second voice: measured, evaluative, impersonal.

The less obvious finding is that readers and machines both reward that convergence, so it reinforces itself. Expert readers couldn't reliably tell LLM-written abstracts from human ones. LLM-edited abstracts got the highest clarity ratings and were preferred 55% of the time even when authorship was disclosed Can readers tell LLM abstracts from human ones?. On the machine side, rewriting a paper's rhetoric while keeping its science unchanged measurably shifts LLM reviewer scores How much does rhetorical style shift AI review scores?. LLM judges also respond to surface cues like formatting and authority signals Can LLM judges be fooled by fake credentials and formatting?. Put these together and authors have a reason to write in the style that machine reviewers and human skimmers both rate well. That narrows style through incentives, not only through copying.

There is a check on the simplest version of this story. The apparent preference of LLM-assisted reviewers for LLM-written papers disappears once paper quality is controlled for Do LLM reviewers actually favor LLM-written papers?. A randomized ICML trial found that banning LLM use, versus allowing limited use, barely changed scores or decisions Does banning LLM use in peer review change review outcomes?. So the evidence points to style converging without clearly changing outcomes. That makes it a quieter problem than bias: reviews may come to sound alike while deciding roughly the same things. To answer whether they actually sound alike, someone would need to measure vocabulary diversity directly, and this corpus doesn't yet contain that work.


Sources 8 notes

How much peer review text shows signs of LLM modification?

Analysis of reviews from ICLR 2024, NeurIPS 2023, CoRL 2023, and EMNLP 2023 estimates this population share using distributional methods rather than per-review classification. Rates were higher in low-confidence, rushed, and less-engaged reviewers.

Does LLM generation explore competing claims while producing text?

Token prediction trains models to continue toward the training distribution, not to explore logically related counterpositions. This smoothness in process produces smooth claims that multiply without generating new perspectives.

Why do LLMs produce such different writing in chat versus posts?

The same model produces sycophantic chat (shaped by RLHF on conversational data) and falsely objective posts (shaped by published prose training). Each register inherits failure modes from its training distribution rather than representing different models or subsystems.

Can readers tell LLM abstracts from human ones?

Readers with ML expertise struggle to identify LLM-generated content reliably, tending to assume human involvement across all abstract types. However, LLM-edited abstracts received highest clarity ratings and were preferred 55% of the time when authorship was disclosed.

How much does rhetorical style shift AI review scores?

Rewriting manuscripts' rhetoric while preserving scientific content moves LLM reviewer scores measurably. Evidence framing and novelty stance produce the largest contrasts; effects vary by reviewer model and interact with the paper's original quality tier.

Show all 8 sources
Can LLM judges be fooled by fake credentials and formatting?

Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.

Do LLM reviewers actually favor LLM-written papers?

Across 125,000+ reviews, the apparent favoritism of LLM-assisted reviewers toward LLM papers disappears once paper quality is held constant. LLM papers cluster among weaker submissions, creating a spurious interaction driven by LLM reviewers' general leniency toward lower-quality work.

Does banning LLM use in peer review change review outcomes?

A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.