Can GPT-4 feedback match what human reviewers catch?
Does an LLM reviewing system raise the same points that human reviewers do? This matters because it tests whether AI could complement or replace human peer review for research papers.
The study reports that GPT-4, run on a paper's PDF through a pipeline that returns structured feedback, raises points that overlap with human reviewers' points about as often as two human reviewers overlap with each other. Across 3,096 papers in 15 Nature family journals, overlap was 30.85% for GPT-4 against an individual reviewer and 28.58% for two reviewers; across 1,709 ICLR papers it was 39.23% against 35.25%. In a prospective user study of 308 researchers from 110 US institutions, 57.4% found the feedback helpful or very helpful and 82.4% found it more beneficial than feedback from at least some human reviewers. These are the study's own measurements.
The overlap is a product of matching, not of a quality judgment. GPT-4 extracts the criticisms in each review, then a second pass pairs them and rates each pair from 5 to 10; only pairs rated 7 ("Strongly Related") or above are kept, because lower ratings "did not always align with human evaluations." Hand checks gave a matching F1 of 0.824 (recall 0.878, precision 0.777), and three annotators agreed on 89.8% of pairs. The measure is therefore agreement about which points get raised, not whether they are right. Overlap was highest for rejected ICLR papers, which the authors read as papers with "more apparent issues or flaws that both human reviewers and LLMs can consistently identify."
Against the nearest notes, the study measures breadth, not reasoning depth. Can structured pipelines make LLM novelty assessment reliable? reports 86.5% alignment with human reviewers, but on one judgment, novelty, split into stages; this excerpt uses a single prompt over the whole manuscript. Its limitations cut against Can inference scaling help reviewers catch errors humans miss?, which spends inference compute to check proofs and experiments line by line: participants say GPT-4 "struggles to provide in-depth critique of method design." The 82.4% figure is a perceived-benefit rating, the kind of signal Does polished writing actually signal better quality work? warns against reading as merit.
The excerpt does not establish accuracy. Overlap with reviewers, who also miss things, is a proxy for usefulness, not a check against errors. It also leaves open how much of each paper the model read. The abstract promises comments on "the full PDFs," but the method passes only "the initial 6,500 tokens" to a model with an 8,192-token limit. The figures disagree internally: the abstract gives 43.80% overlap for rejected ICLR papers, while the discussion gives that figure for two human reviewers and 47.09% for GPT-4 against humans, so the abstract appears to report the human-human number. The excerpt also omits how the user study measured benefit relative to human reviewers and how respondents were recruited. The supportable implication is narrower than the headline: LLM feedback can match reviewers' point coverage and be perceived as useful, which makes it a complement to review rather than a substitute for methodological critique.
Inquiring lines that read this note 8
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can AI systems perform peer review as effectively as humans?- Can LLM reviewers catch technical issues that human reviewers miss?
- How did researchers measure whether GPT-4 and human reviewers identified the same issues?
- Does matching reviewer points actually mean the feedback is accurate or correct?
- Why did rejected papers show higher overlap between GPT-4 and human reviewers?
- Could AI feedback work as a substitute for human peer review entirely?
- Does an automated reviewer's output actually match human review accuracy?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can structured pipelines make LLM novelty assessment reliable?
Explores whether breaking novelty assessment into extraction, retrieval, and comparison stages helps LLMs align with human peer reviewers and produce more rigorous, evidence-based evaluations.
contrasts a three-stage novelty pipeline's alignment score with this study's single-pass overlap measure
-
Can inference scaling help reviewers catch errors humans miss?
Explores whether spending extra compute at review time—checking proofs and experiments line by line—can surface deep flaws that evade human expert reviewers, and how this scales with AI-assisted submissions.
contrasts: this study's participants say single-pass GPT-4 struggles with in-depth method critique
-
Does polished writing actually signal better quality work?
When evaluators judge applications and manuscripts, does rhetorical sophistication predict merit, or does it distract from verifiable evidence of competence and rigor?
the 82.4% helpfulness rating is perceived benefit, the signal that review cautions against treating as merit
-
Did conference reviewers prefer AI reviews over human ones?
AAAI-26 ran a live pilot adding labeled AI reviews to 22,977 papers alongside human reviews. A survey asked participants which reviews were more useful, particularly on technical accuracy and research suggestions.
Evidence for: a conference-wide pilot survey also reports participants preferred AI reviews to human ones on technical accuracy
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Can large language models provide useful feedback on research papers? A large-scale empirical analysis
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- Stop Automating Peer Review Without Rigorous Evaluation
- Pangram Predicts 21% of ICLR Reviews are AI-Generated
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- Can Large Language Models Make the Grade? An Empirical Study Evaluating LLMs Ability to Mark Short Answer Questions in K-12 Education
Original note title
GPT-4 feedback overlaps with one human reviewer about as much as two reviewers overlap with each other — 30.85 percent against 28.58 percent