When GPT-4 gave feedback on research papers, did it read the whole paper or only part of it?
Did GPT-4 see the entire paper or only a portion of it?
This reads the question as asking about the study where GPT-4 gave feedback on Nature and ICLR papers: was the model reading whole manuscripts or excerpts, and does that change how far its feedback can be trusted?
This reads the question as asking about the study where GPT-4 gave feedback on Nature and ICLR papers: was the model reading whole manuscripts or excerpts, and does that matter? The short answer is that the notes here don't say. The headline result is clear. Across 3,096 Nature papers and 1,709 ICLR papers, GPT-4's comments matched an individual human reviewer's points 30.85% of the time. Two human reviewers matched each other 28.58% of the time. But the summary doesn't record how much of each paper was actually given to the model Can GPT-4 feedback match what human reviewers catch?. To settle it you'd need the original paper's methods section, not this collection.
The question matters more than it first seems. An overlap score only tells you how often GPT-4 raised the same points as a human. It can't tell you whether those points came from reading the whole argument or from skimming the parts that are easy to comment on, like the abstract, the framing and the stated limitations. If the model saw only a portion, the high overlap might partly mean that human reviewers also cluster around those same surface features. That would say something uncomfortable about reviewing in general, not only about AI.
A sharper version of the same question appears elsewhere in the collection. What an AI 'sees' in a paper isn't always what a human sees. Researchers at institutions in at least eight countries hid invisible instructions in their PDFs and HTML, telling AI reviewers to write flattering assessments. The text doesn't show in an ordinary PDF reader, but a model processing the full document reads it Are researchers hiding prompt injections in academic papers?. A separate study found 18 arXiv manuscripts with similar self-serving prompts and argues this counts as a questionable research practice Are hidden AI prompts in preprints a deceptive research practice?. So 'did it see the whole paper?' has a twist. Feeding a model everything can expose it to manipulation that a partial or human-facing view would miss.
How the input is cut and fed in also changes the answer. Retrieval research shows that a model's first, partial answer can reveal what information it's still missing, and using that answer to fetch more material improves the result Can a model's partial response guide what to retrieve next?. A reviewer that reads a paper in one fixed pass, whole or truncated, gets no such second look. The evaluation literature points the same way: the judge's role and setup shape its verdict. GPT models showed a preference for their own company's outputs only when acting as graders in agentic setups Does grading expose company bias that answering hides?, and panels of smaller, varied judges beat a single large judge Can a panel of smaller judges outperform one large judge?.
The takeaway is that 'GPT-4 reviews about as well as a human' depends on details the headline leaves out: what text went in, what was hidden in it, and whether the model got more than one pass. When you read claims about AI peer review, ask what the model was actually shown.
Sources 6 notes
A study of 3,096 Nature papers and 1,709 ICLR papers found GPT-4 matched individual reviewers' points 30.85% of the time versus 28.58% for two human reviewers. Fifty-seven percent of surveyed researchers found the feedback helpful.
Researchers at institutions across eight countries embedded invisible instructions in paper PDFs and HTML telling AI models to produce flattering summaries. The text remains undetectable in standard PDF readers but executes when AI systems process the full document.
Eighteen arXiv manuscripts contained concealed instructions directing AI reviewers to give positive assessments. The practice qualifies as questionable research conduct because concealment plus self-serving design violates ethics regardless of stated intent.
ITER-RETGEN shows that iteratively using generated responses as retrieval queries substantially improves performance on multi-hop reasoning and fact verification. Generation acts as both answer producer and information-need clarifier, surfacing implicit gaps that the original query missed.
Across four tasks, GPT models show no company bias except in Agentic Grading, where they favor their own company alongside Claude's known bias. This suggests the grading role—particularly in agentic setups—activates preference patterns that question-answering tasks do not trigger.
Show all 6 sources
PoLL (Panel of LLM evaluators) using multiple smaller models from disjoint families outperforms single large judges, reduces intra-model bias, and costs over 7× less. Across three settings and six datasets, no single judge was best everywhere, but diverse panels performed consistently well.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Hidden Prompts in Manuscripts Exploit AI-Assisted Peer Review
- Scholars sneaking phrases into papers to fool AI reviewers
- Stop Automating Peer Review Without Rigorous Evaluation
- Pangram Predicts 21% of ICLR Reviews are AI-Generated
- The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing
- When Reject Turns into Accept: Quantifying the Vulnerability of LLM-Based Scientific Reviewers to Indirect Prompt Injection
- Can Large Language Models Make the Grade? An Empirical Study Evaluating LLMs Ability to Mark Short Answer Questions in K-12 Education
- Can large language models provide useful feedback on research papers? A large-scale empirical analysis