SYNTHESIS NOTE
Topics›Domain Specialization›this note

Can two-stage review and badges fix AI conference peer review?

A position paper diagnoses AI conference review failures across authors, reviewers, and venues, proposing staged author feedback on reviews and a reviewer reward system. Does this approach actually reduce bias and improve review quality?

Synthesis note · 2026-10-06 · sourced from Domain Specialization

The paper's authors argue that the peer review crisis at major AI conferences cannot be placed on reviewers alone. They name three parties, authors, reviewers and "the System" (the venue and OpenReview), and say all three "share responsibility for the current problems." Author misconduct, they add, "can only be addressed through policy enforcement and detection tools," so they focus on reviewer accountability: a two-stage bi-directional review in which authors rate reviews, and a systematic reviewer reward system.

The first mechanism changes the order of release. Today all reviews and ratings reach authors at once. Under the proposal, each review's summary, strengths and clarifying questions come first, and authors grade them on the reviewer's comprehension and the constructiveness of the questions. Weaknesses and ratings follow, and authors cannot revise their evaluations afterward. The authors argue that rating a review before seeing its verdict blocks "retaliatory scoring." An LLM review, visible only to authors, is meant to be "a psychological deterrent" for reviewers and "a soft reference point" for spotting machine-written reviews. For rewards, reviewers would earn verifiable badges. The paper concedes that badges could reward lenient reviewing, and asks for a metric that rewards thoroughness over positivity.

The diagnosis rests on NeurIPS studies, as the paper reports them. In consistency experiments in 2014 and 2021, "16-23% of papers could have been either accepted or rejected based on the reviewing group." A 2022 NeurIPS study of review quality found "author-outcome bias" and "elongated review bias," the latter rating longer reviews higher "even when the two reviews contain the same information." The two-stage design targets the first, and capping the length of stage-one content targets the second. Against the nearest notes, this is a different lever from the one the ICML 2026 trial tested. Does banning LLM use in peer review change review outcomes? found rules barely moved outcomes, while this paper bets on visibility, which it does not test. Its LLM review is a reference for authors, not a reviewer, unlike Can inference scaling help reviewers catch errors humans miss?, where a model checks proofs and experiments line by line. The length bias also parallels the surface sensitivity in How much does rhetorical style shift AI review scores?, here measured in human reviewers.

The excerpt does not establish that the fix works. The paper reports no pilot. Its Section 5.1 asks for a large survey and a small-track pilot first, and Section 5.2 names the main obstacle as getting venues and OpenReview to build the feedback and reward systems, plus cost: ICML 2024 cut in-kind compensation for its top 10% of reviewers to a smaller pool. The claim that stage-one release "could prevent retaliatory scoring" is an inference from the bias finding, not a measured effect, and the NeurIPS results come without effect sizes. The 10,000-submissions figure carries no citation in the excerpt. The strongest support is for the diagnosis, that review is noisy and rating-biased. The prescription is a plausible design that a venue could test as a pilot before adopting it as policy.

Inquiring lines that read this note 46

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can AI systems perform peer review as effectively as humans? Do restrictions on reviewer LLM use actually shape peer review behavior? How can we detect and account for LLM involvement in academic writing? What explains the gap between benchmark scores and true reasoning capability?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 61 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

AI conference review problems are shared by three parties, so the position argues for two-way author feedback and reviewer rewards