Can AI reviewers catch the broken proofs and shaky experiments human peer reviewers let slip through, if they're set up to dig?
Can LLM reviewers catch technical issues that human reviewers miss?
This explores whether AI reviewers can find real flaws in research papers, like broken proofs or shaky experiments, that human peer reviewers miss, and what kind of AI review actually manages it.
This explores whether AI reviewers can find real technical flaws in papers, such as broken proofs or shaky experiments, that human reviewers let through. The short answer is yes, but only when the AI is set up to dig. The strongest evidence comes from an agentic reviewer that spends extra computing time checking proofs and experiments line by line. It found math errors noticeably better than a model asked to review in a single pass, and it flagged critical flaws in papers that had already been accepted at top venues like STOC and ICML Can inference scaling help reviewers catch errors humans miss?. What made the difference was not a smarter model reading the paper once. It was a model given the time and structure to check each step methodically, which is the kind of tedious verification that busy human reviewers rarely have time for.
The same pattern shows up in novelty assessment, which is judging whether a paper's contribution is actually new. When an LLM was split into stages (pull out the paper's claims, find related work, then compare), its reasoning matched human reviewers' reasoning 86% of the time, and it beat models asked for a single overall judgment Can structured pipelines make LLM novelty assessment reliable?. Plain, unstructured AI feedback looks more like another ordinary reviewer. GPT-4's comments overlapped with any one human reviewer's points about as often as two human reviewers overlapped with each other Can GPT-4 feedback match what human reviewers catch?. That finding cuts both ways: the AI catches some things a given human misses, but mostly because human reviewers already disagree a lot about what matters.
There are sharp limits. LLM judges can be fooled by fake credentials and polished formatting, and these tricks work no matter what the paper actually says Can LLM judges be fooled by fake credentials and formatting?. In a study of more than 125,000 reviews, LLM-assisted reviewers seemed to favor LLM-written papers. That turned out to be an illusion: those reviewers were simply more lenient toward weaker work in general, and LLM-written papers tended to be weaker Do LLM reviewers actually favor LLM-written papers?. Leniency toward weak work is the opposite of catching what humans miss. A further worry is that frontier models tend to fail through subtle errors that leave the surface looking intact, rather than through obvious gaps Does model capability change how documents degrade?. An AI review that sounds thorough can therefore be wrong in ways that are hard to spot.
The most practical role so far is making human reviewers better rather than replacing them. At ICLR 2025, a randomized trial gave reviewers optional feedback on their own reviews from Claude-based agents. Over a quarter of reviewers revised their reviews, and blinded raters judged the revised versions more specific and informative Can LLM feedback help peer reviewers improve their own reviews?. Conferences are also learning where AI is easiest to police. ICLR 2026 treated AI-detector flags as one input for human area chairs to weigh, but desk-rejected papers with confirmed fabricated references, because a fake citation is something anyone can verify How can conferences detect and handle LLM misuse in peer review?. Meanwhile, an ICML 2026 experiment found that banning LLM use versus allowing limited use barely changed scores or decisions, and many reviewers broke whichever rule they were given Does banning LLM use in peer review change review outcomes?.
The point you might not expect: human reviewers often miss flaws not because they lack the skill but because they can't bring the relevant reasoning to mind while reviewing. Experiments with 640 employees found that people caught more errors when the reasoning needed for checking was easy to call up at that moment, for example through self-generated explanations or retrieval cues Can reviewers access what they know when checking LLM outputs?. That suggests a useful split of the work. The AI's main advantage is patient, step-by-step checking, and the human's main weakness is having the right knowledge at hand when it counts. AI reviewers help most when they put the specific claims and checks in front of a human, not when they hand over a verdict.
Sources 10 notes
PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.
A three-stage pipeline (extract claims, retrieve related work, compare) reached 86.5% reasoning alignment and 75.3% conclusion agreement with human reviewers on 182 ICLR submissions, outperforming holistic LLM baselines.
A study of 3,096 Nature papers and 1,709 ICLR papers found GPT-4 matched individual reviewers' points 30.85% of the time versus 28.58% for two human reviewers. Fifty-seven percent of surveyed researchers found the feedback helpful.
Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.
Across 125,000+ reviews, the apparent favoritism of LLM-assisted reviewers toward LLM papers disappears once paper quality is held constant. LLM papers cluster among weaker submissions, creating a spurious interaction driven by LLM reviewers' general leniency toward lower-quality work.
Show all 10 sources
DELEGATE-52 shows weaker LLMs degrade documents through visible deletion, while frontier models degrade through subtle corruption that preserves surface integrity. This shift makes frontier failures harder to detect and potentially more dangerous at workflow scale.
A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.
Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.
A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.
Two experiments with 640 employees showed that error detection improved when verification-relevant reasoning was accessible at review time. Self-generated explanations and retrieval cues strengthened detection, revealing a third failure mode beyond capability or engagement gaps.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- Stop Automating Peer Review Without Rigorous Evaluation
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026
- Can large language models provide useful feedback on research papers? A large-scale empirical analysis
- Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot