Why did GPT-4's reviews line up with human reviewers' comments more often on rejected papers than on accepted ones?
Why did rejected papers show higher overlap between GPT-4 and human reviewers?
This explores why GPT-4's review comments lined up more closely with human reviewers' comments on papers that were ultimately rejected than on papers that were accepted, and what that tells us about what AI reviewers actually see.
This explores why GPT-4's review feedback matched human reviewers' points more often on rejected papers than on accepted ones. The note in this collection covers the headline result, not the accept/reject breakdown. Across thousands of Nature and ICLR papers, GPT-4 matched an individual human reviewer's points about as often as two humans matched each other (30.85% vs 28.58%) Can GPT-4 feedback match what human reviewers catch?. The original study did report the split you're asking about. Its likely reading is that weaker papers have obvious, widely visible problems: missing baselines, thin experiments, unclear claims. Those are exactly the issues that a careful generalist, human or model, will flag. Accepted papers have fewer surface flaws, so the remaining criticism is more specific and depends more on each reviewer's own judgment. That leaves less common ground for any two reviewers to share.
Other notes in the collection support this reading. AI reviewers show a 'hivemind effect': they agree with each other more than humans do, which suggests they converge on the same generic, easy-to-spot concerns Can AI systems safely replace human peer reviewers?. Generic concerns are the kind that pile up on weak papers. So the higher overlap on rejected papers may say less about GPT-4's insight and more about the kind of flaw that's easy to see.
A separate large study shows the same weak-paper effect causing a misleading result. LLM-assisted reviewers seemed to favor LLM-written papers. The effect disappeared once paper quality was held constant, because LLM papers clustered among weaker submissions Do LLM reviewers actually favor LLM-written papers?. The lesson carries over: before reading meaning into any AI-versus-human review pattern, check whether paper quality is the real driver. At Agents4Science, accepted papers also involved more human guidance than rejected ones Do accepted papers need more human guidance than rejected ones?. That's another hint that 'rejected' often marks a recognizable kind of weakness.
Here's the twist. Matching human reviewers on obvious flaws is not the same as catching the flaws humans miss. An agentic reviewer that spends extra compute checking proofs and experiments line by line found critical errors in papers that had already passed human review at STOC and ICML Can inference scaling help reviewers catch errors humans miss?. High overlap on rejected papers shows that GPT-4 can reproduce the consensus critique. The more valuable target is the strong-looking paper with a hidden flaw, where overlap with humans would be low by design.
The practical use follows from that. AI feedback seems most dependable as a first pass that catches the obvious problems, which is why it can help authors early and help reviewers sharpen their comments Can LLM feedback help peer reviewers improve their own reviews?. It's least dependable for the close calls at the acceptance line.
Sources 6 notes
A study of 3,096 Nature papers and 1,709 ICLR papers found GPT-4 matched individual reviewers' points 30.85% of the time versus 28.58% for two human reviewers. Fifty-seven percent of surveyed researchers found the feedback helpful.
AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.
Across 125,000+ reviews, the apparent favoritism of LLM-assisted reviewers toward LLM papers disappears once paper quality is held constant. LLM papers cluster among weaker submissions, creating a spurious interaction driven by LLM reviewers' general leniency toward lower-quality work.
At Agents4Science, organizers observed that accepted papers carried more human input than rejected ones, with humans concentrated in design and hypothesis work while AI gained autonomy in analysis and writing. The pattern emerged from self-reported disclosure tiers across four research stages.
PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.
Show all 6 sources
A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
- Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?
- Can large language models provide useful feedback on research papers? A large-scale empirical analysis