INQUIRING LINE

When only a paper's wording changes, with the science untouched, one study saw its AI review score rise nearly half a point.

Do AI-generated research reviews score papers higher than human reviewers do?

This explores whether AI reviewers grade research papers more generously than human reviewers do. The corpus doesn't measure that score gap directly, but it has several nearby findings on how AI judgment of papers differs from human judgment.


This explores whether AI reviewers are more generous graders than human reviewers. The short answer: none of these notes runs a direct comparison of AI scores against human scores on the same papers, so the collection can't say whether AI reviewers are harsher or kinder on average. It does show something that may matter more. The problem with AI reviewers isn't how high or low their scores are. It's what moves those scores. One study found that simply rewriting a paper's text, with no change to the science, raised AI review scores by about 0.45 points. The same study found that AI reviewers agree with each other far more than human reviewers do, which the authors call a 'hivemind' Can AI systems safely replace human peer reviewers?. So a paper can earn a higher AI score by being polished for the reviewer, not by being better. And because AI reviewers share the same blind spots, a group of them doesn't give you the range of opinions you'd get from several humans.

The question can also be turned around: how do human reviewers score papers written by AI? Two notes describe the same experiment. A fully AI-generated paper averaged 6.33 from human reviewers at an ICLR 2025 workshop, enough to be accepted and in the top 45% of submissions Can AI systems generate research papers that pass peer review?. Afterward, the authors found a citation error and decided that none of their three submissions was good enough for the main conference Can AI-generated papers pass peer review undetected?. The papers' own authors judged them more harshly than the reviewers who passed one of them. Human reviewers can be fooled too, and that's one reason the score question is hard to settle.

Scoring isn't the only thing a reviewer does. An agentic reviewer called PAT spends extra computing time checking proofs and experiments line by line. It found serious flaws in papers that had already passed human review at STOC and ICML Can inference scaling help reviewers catch errors humans miss?. At ICLR 2025, AI feedback on human reviewers' drafts led 27% of them to revise, and blinded raters judged the revised reviews more informative Can LLM feedback help peer reviewers improve their own reviews?. A randomized trial at ICML 2026 found that banning LLM use versus allowing limited use barely changed scores or decisions, and many reviewers broke whichever rule they were given Does banning LLM use in peer review change review outcomes?. At least in how reviewers actually use AI assistance today, it doesn't appear to push scores up or down.

A bigger-picture view: one survey of 230 publications describes AI paper writing and AI reviewing as an arms race. AI lets researchers produce more papers, venues automate evaluation in response, authors learn to game the automated reviewers, and defenses follow Does AI create a coupled arms race in research production and review?. In that setting, any score gap between AI and human reviewers would keep changing as authors adapt. Some proposals build defenses in from the start. aiXiv runs repeated rounds of automated review and revision with protections against prompt injection Can automated review loops handle AI-generated research at scale?. Another proposal attacks human scoring biases, such as ratings that track review length, by having authors rate review quality before they see the verdict Can two-stage review and badges fix AI conference peer review?.

The takeaway you may not have expected: whether AI reviewers score higher on average matters less than what moves their scores. An AI reviewer that rewards polish can be gamed, and many AI reviewers that agree with each other remove the disagreement that makes peer review work. A direct score-calibration study would be a worthwhile addition to this collection.


Sources 9 notes

Can AI systems safely replace human peer reviewers?

AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.

Can AI systems generate research papers that pass peer review?

AI Scientist-v2 submitted three fully autonomous manuscripts to ICLR; one averaged 6.33 from reviewers and ranked in the top 45% of workshop submissions. The authors acknowledged the work does not yet meet top-tier conference standards and withdrew the accepted paper before publication.

Can AI-generated papers pass peer review undetected?

Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.

Can inference scaling help reviewers catch errors humans miss?

PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.

Can LLM feedback help peer reviewers improve their own reviews?

A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.

Show all 9 sources
Does banning LLM use in peer review change review outcomes?

A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.

Does AI create a coupled arms race in research production and review?

A survey of 230 publications reveals production scaling, evaluation automation, manipulation, defenses, evasion, and ecosystem feedback as linked response relations among actors. Evidence is strongest for early stages and weakens toward long-horizon adaptation and feedback.

Can automated review loops handle AI-generated research at scale?

aiXiv demonstrates that iterative review-refine cycles with automated retrieval-augmented evaluation and prompt-injection defenses measurably enhance proposal and paper quality, addressing the structural gap where AI-generated research lacks appropriate publication venues.

Can two-stage review and badges fix AI conference peer review?

Authors, reviewers, and venues all contribute to peer review failures at major AI conferences. A proposed two-stage system lets authors rate review quality before seeing verdicts, and a badge system rewards reviewer thoroughness, targeting measured biases like rating-length correlation.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.