Reviewers who are rushed or unsure lean harder on AI, but the evidence doesn't directly compare them with the editors who weigh their reviews.
Do metareviewers and regular reviewers use LLMs differently in peer review?
This explores whether the people who write individual paper reviews and the people who weigh those reviews to make a final call (metareviewers or area chairs) use LLMs in different ways, and what the collection says about each role.
This explores whether the people who write individual reviews and the metareviewers or area chairs who combine them into a decision use LLMs differently. The short answer: the collection does not have a study that compares the two roles directly. What it does show is that the two roles meet LLMs from opposite sides. Reviewers are mostly studied as users of LLMs. Metareviewers show up mostly as the people who judge that use.
On the reviewer side, the evidence is about who leans on LLMs and what changes when they do. One estimate puts substantially LLM-modified text in 6.5 to 16.9 percent of reviews at recent AI conferences. Those rates were higher among reviewers who reported low confidence, submitted close to the deadline, or engaged less with the authors' responses How much peer review text shows signs of LLM modification?. So LLM use tracks how pressed or unsure a reviewer is, not just the reviewer's role. When a conference randomly assigned reviewers to a ban or to limited use, scores and decisions barely moved. Many reviewers broke whichever rule they were given Does banning LLM use in peer review change review outcomes?. LLMs can also help reviewers, not just stand in for them. At ICLR 2025, optional feedback from an AI agent led 27 percent of reviewers to revise, and their revised reviews were rated more specific Can LLM feedback help peer reviewers improve their own reviews?.
Metareviewers appear in a different position. At ICLR 2026, program chairs did not use LLM detectors as automatic filters. They sent detector flags to area chairs as one piece of evidence among several, so a human made the call on possible misuse. The only automatic penalty was desk rejection for confirmed fabricated references How can conferences detect and handle LLM misuse in peer review?. In other words, the metareviewer's job is shifting toward weighing how trustworthy each review is, partly by judging whether a machine wrote it.
This matters because of what LLM-written reviews tend to look like. AI reviewers agree with each other more than human reviewers do, and rewriting a paper's text can raise their scores without improving the science Can AI systems safely replace human peer reviewers?. Reviewers who use LLMs also tend to go easy on weaker papers. That leniency is why they appear to favor LLM-written papers, since those papers cluster among weaker submissions Do LLM reviewers actually favor LLM-written papers?. A metareviewer who reads several LLM-assisted reviews may therefore see agreement that comes from shared machine habits rather than from independent judgments. Spotting this is hard: even ML experts cannot reliably tell LLM text from human text Can readers tell LLM abstracts from human ones?. Research on human oversight of AI adds that checkers catch more errors when the reasoning they need is easy to reach at the moment they review Can reviewers access what they know when checking LLM outputs?.
The takeaway you might not expect: the open question is less how metareviewers use LLMs and more how they should discount reviews shaped by them. The collection has nothing yet on metareviewers writing their own summaries with LLMs. If you want to follow this thread, start with the ICLR 2026 note.
Sources 8 notes
Analysis of reviews from ICLR 2024, NeurIPS 2023, CoRL 2023, and EMNLP 2023 estimates this population share using distributional methods rather than per-review classification. Rates were higher in low-confidence, rushed, and less-engaged reviewers.
A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.
A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.
Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.
AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.
Show all 8 sources
Across 125,000+ reviews, the apparent favoritism of LLM-assisted reviewers toward LLM papers disappears once paper quality is held constant. LLM papers cluster among weaker submissions, creating a spurious interaction driven by LLM reviewers' general leniency toward lower-quality work.
Readers with ML expertise struggle to identify LLM-generated content reliably, tending to assume human involvement across all abstract types. However, LLM-edited abstracts received highest clarity ratings and were preferred 55% of the time when authorship was disclosed.
Two experiments with 640 employees showed that error detection improved when verification-relevant reasoning was accessible at review time. Self-generated explanations and retrieval cues strengthened detection, revealing a third failure mode beyond capability or engagement gaps.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026
- Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on ArXiv