When an AI reviewer's points match a human's, is that proof it's right, or just agreement with one reviewer's partial view?
Does matching reviewer points actually mean the feedback is accurate or correct?
This explores whether an AI reviewer that raises the same points as human peer reviewers is actually producing correct feedback, or whether agreement and accuracy are two different things.
This asks whether an AI reviewer that 'matches' human reviewers is actually right, or only agreeing. The corpus points clearly toward the second: overlap measures agreement, not correctness, and the two can come apart in surprising ways. The best-known result here is that GPT-4's feedback overlapped with an individual human reviewer's points 30.85% of the time, while two human reviewers overlapped with each other only 28.58% of the time Can GPT-4 feedback match what human reviewers catch?. That sounds like parity with humans. But look at what the baseline shows: human reviewers of the same paper agree on only about 30% of their points. Matching a reviewer means matching one partial, idiosyncratic view of the paper, not some agreed-upon truth.
The less obvious problem is that the most valuable feedback may be the feedback that matches nobody. An agentic reviewer that spends extra compute checking proofs and experiments line by line found critical flaws in papers that had already passed human review at STOC and ICML Can inference scaling help reviewers catch errors humans miss?. Scored by overlap with the human reviews, those catches would count as misses, because no human raised them. A 'match the humans' metric rewards an AI for repeating what reviewers already noticed and penalizes it for finding what they missed.
Agreement can also be produced without any accuracy behind it. AI reviewers show a 'hivemind' effect: they agree with each other more than humans do. And a zero-shot rewrite of a paper's text raised AI review scores by 0.45 points without changing the science at all Can AI systems safely replace human peer reviewers?. Consensus among AI systems therefore says little, and their judgments respond to surface polish. Consumer ratings research shows a similar pattern outside science: later ratings drift toward earlier ones, so a high-agreement aggregate can partly reflect herding rather than independent judgments of quality Do online ratings actually reflect independent customer opinions?. A related agent-safety note makes the general point: a scoring function can compute perfectly and still report something misleading if what it measures has been shaped upstream Can a correct scoring function still mislead about task performance?.
So what would test whether feedback is correct? The corpus offers a few alternatives to overlap. One is to verify specific claims directly, as the proof-checking reviewer does. Another is to have blinded raters judge whether a revised review is more specific and informative. That was the measure used in the ICLR 2025 trial, where 27% of reviewers updated their reviews after LLM feedback Can LLM feedback help peer reviewers improve their own reviews?. A third is to let authors rate how good a review was before they see the accept/reject decision, so their judgment isn't colored by the outcome Can two-stage review and badges fix AI conference peer review?. Each of these asks 'was this right or useful?' instead of 'did it sound like the others?' The takeaway: when a headline says an AI reviewer 'matches humans,' check what humans match each other on first. The bar is lower, and less tied to correctness, than it sounds.
Sources 7 notes
A study of 3,096 Nature papers and 1,709 ICLR papers found GPT-4 matched individual reviewers' points 30.85% of the time versus 28.58% for two human reviewers. Fifty-seven percent of surveyed researchers found the feedback helpful.
PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.
AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.
Moe and Trusov decomposed ratings into baseline quality, social-dynamics influence, and error, finding that prior ratings meaningfully affect subsequent ones. These effects have both immediate sales impact and long-term compounding effects through future ratings, though high opinion variance can eventually dampen the distortion.
A scoring function can compute correctly over inputs while still attesting to the wrong thing if an agent has altered those inputs or their provenance outside the intended task path. Verification of the function itself is necessary but insufficient in stateful systems.
Show all 7 sources
A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.
Authors, reviewers, and venues all contribute to peer review failures at major AI conferences. A proposed two-stage system lets authors rate review quality before seeing verdicts, and a badge system rewards reviewer thoroughness, targeting measured biases like rating-length correlation.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?
- Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
- Can large language models provide useful feedback on research papers? A large-scale empirical analysis
- How to Find Fantastic AI Papers: Self-Rankings as a Powerful Predictor of Scientific Impact Beyond Peer Review
- The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing