When AI writes an entire peer review, its scores bunch up in the middle. Is it going easy on weak papers?
What mechanisms drive rating compression in fully LLM-generated peer reviews?
This explores why reviews written entirely by an LLM tend to bunch their scores into a narrow band, rarely calling a paper terrible or brilliant, and what in the models or the review setup causes that bunching.
This explores why fully LLM-written reviews squeeze their scores toward the middle instead of spreading papers out the way human reviewers do. First, a limit on the evidence: no note in this collection measures rating compression directly. What the collection does have is several findings that each explain part of it, and together they make a fairly clear picture.
The first mechanism is leniency at the bottom of the scale. A study of more than 125,000 reviews set out to test whether LLM-assisted reviewers favor LLM-written papers. That apparent favoritism disappeared once paper quality was taken into account. What remained was simpler: LLM reviewers are generally more lenient toward weaker work Do LLM reviewers actually favor LLM-written papers?. If weak papers get pulled up while strong papers can't go much higher, the spread of scores shrinks from below. Compression doesn't need a deliberate bias toward the middle. It can be a floor that rises unevenly.
The second mechanism is sameness across reviewers. AI reviewers show a 'hivemind' effect: across papers they agree with each other more than human reviewers do Can AI systems safely replace human peer reviewers?. Peer review uses several reviewers because their disagreements carry information. When every reviewer is drawn from roughly the same model, their errors and their caution point the same way, so averaging their scores adds little new information. A related finding helps explain why: models favor text they can recognize as their own, and that preference grows in step with how well they recognize it Do LLMs favor their own text because they recognize it?. A model reviewing polished, model-like prose may see little to object to.
The third mechanism is that surface features move scores more than substance does. Zero-shot rewrites of a paper's text raised AI review scores by 0.45 points without changing the science Can AI systems safely replace human peer reviewers?. LLM judges also fall for fake authority signals and rich formatting Can LLM judges be fooled by fake credentials and formatting?. If presentation drives the score, then papers of very different quality that are all competently written will get similar ratings. People show the same weakness, which suggests the problem isn't unique to machines: users prefer answers with more citations even when the citations are irrelevant Do users trust citations more when there are simply more of them?.
One connection is speculative but worth following. Measured with rate-distortion theory, LLMs keep the broad shape of a category but drop the fine distinctions that humans preserve Do LLMs compress concepts more aggressively than humans do?. A compressed rating scale could be the same habit applied to judgment: 'competent ML paper' becomes one category, and the differences inside it disappear. The collection also hints at a way out. Reviews improve when the task is broken into steps instead of asking for one overall judgment. A pipeline that extracts claims, retrieves related work and then compares reached 86% agreement with human reviewers on whether papers were novel Can structured pipelines make LLM novelty assessment reliable?. An agentic reviewer that checks proofs line by line found flaws that human experts missed Can inference scaling help reviewers catch errors humans miss?. Checking specific claims gives a reviewer concrete reasons to score a paper low, which a single overall impression rarely does.
Sources 8 notes
Across 125,000+ reviews, the apparent favoritism of LLM-assisted reviewers toward LLM papers disappears once paper quality is held constant. LLM papers cluster among weaker submissions, creating a spurious interaction driven by LLM reviewers' general leniency toward lower-quality work.
AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.
Fine-tuning LLMs to recognize their own summaries increased their preference for those summaries in a linear relationship, suggesting recognition capability drives self-preference bias. The authors present this as initial causal evidence, not proof.
Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.
Analysis of 24,000 Search Arena interactions shows irrelevant citations boost user preference (β=0.273) nearly as much as relevant citations (β=0.285), indicating citation count functions as a decoupled trust heuristic.
Show all 8 sources
Using Rate-Distortion Theory on cognitive datasets, LLMs capture broad category structure but lose fine-grained distinctions humans preserve. LLMs maximize compression efficiency; humans trade compression for contextual meaning that enables situated action.
A three-stage pipeline (extract claims, retrieve related work, compare) reached 86.5% reasoning alignment and 75.3% conclusion agreement with human reviewers on 182 ICLR submissions, outperforming holistic LLM baselines.
PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
- References Improve LLM Alignment in Non-Verifiable Domains