INQUIRING LINE

Do AI conference scores track how much reviewers write, rather than how good the paper is?

Does review length bias affect acceptance decisions at major conferences?

This explores whether the length of a peer review, meaning how much a reviewer writes rather than how good the paper is, is tied to the scores papers get at major AI conferences, and what can be done about it.


This explores whether how much a reviewer writes is tied to the score a paper receives at large AI conferences, separately from how good the paper actually is. The collection has only one source that measures this directly, so a full answer isn't possible here. A position paper on AI conference peer review lists a correlation between review ratings and review length among the biases it measured Can two-stage review and badges fix AI conference peer review?. Its proposed fixes are aimed at problems like that one. In a two-stage process, authors grade how useful a review was before they see the accept/reject decision, so their grade isn't colored by the outcome. A badge system rewards reviewers for thoroughness rather than for word count. The summary doesn't say which way the correlation runs or how much it moves final decisions, so the collection can't yet answer the 'acceptance decisions' part of the question.

A correlation alone doesn't prove bias, and another study shows why. Across more than 125,000 reviews, reviewers who used LLMs seemed to favor LLM-written papers. The effect disappeared once paper quality was taken into account: LLM-written papers were mostly weaker submissions, and LLM-assisted reviewers were more lenient toward weaker work in general Do LLM reviewers actually favor LLM-written papers?. Length could work the same way. A reviewer facing a weak paper may simply have more to say. A real bias would mean length moves scores even when the paper's quality is the same.

The idea that surface features stand in for substance has stronger evidence in nearby areas. When people compare AI answers, irrelevant citations raise their preference almost as much as relevant ones, so the number of citations works as a trust signal on its own Do users trust citations more when there are simply more of them?. AI reviewers are even easier to sway. Rewriting a paper's text without changing its science raised AI review scores by about 0.45 points, and AI reviewers agree with each other more than human reviewers do Can AI systems safely replace human peer reviewers?. If reviews are judged by how long or polished they look, the same kind of shortcut is likely.

Some studies change how reviews are written without looking at their length. At ICLR 2025, AI feedback led 27 percent of reviewers to revise, and independent raters who didn't know which reviews had been revised judged the new versions more specific and informative Can LLM feedback help peer reviewers improve their own reviews?. At ICML 2026, banning LLM use versus allowing limited use barely changed scores or decisions Does banning LLM use in peer review change review outcomes?. Together they suggest that changing how a review reads doesn't necessarily change the verdict. That would be the test case for length bias, but nobody in the collection has measured it.


Sources 6 notes

Can two-stage review and badges fix AI conference peer review?

Authors, reviewers, and venues all contribute to peer review failures at major AI conferences. A proposed two-stage system lets authors rate review quality before seeing verdicts, and a badge system rewards reviewer thoroughness, targeting measured biases like rating-length correlation.

Do LLM reviewers actually favor LLM-written papers?

Across 125,000+ reviews, the apparent favoritism of LLM-assisted reviewers toward LLM papers disappears once paper quality is held constant. LLM papers cluster among weaker submissions, creating a spurious interaction driven by LLM reviewers' general leniency toward lower-quality work.

Do users trust citations more when there are simply more of them?

Analysis of 24,000 Search Arena interactions shows irrelevant citations boost user preference (β=0.273) nearly as much as relevant citations (β=0.285), indicating citation count functions as a decoupled trust heuristic.

Can AI systems safely replace human peer reviewers?

AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.

Can LLM feedback help peer reviewers improve their own reviews?

A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.

Show all 6 sources
Does banning LLM use in peer review change review outcomes?

A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.