Can conference organizers keep bad papers out, or can they only catch the problems they can actually check?
What role do conference organizers play in accepting problematic articles?
This explores how conference organizers (program chairs, editors, venue designers) let flawed, fraudulent or AI-gamed papers through, and what tools they actually have to stop them.
This explores the gatekeeping role of conference and journal organizers: where they contribute to bad papers getting accepted, and where their influence works or falls short. The corpus suggests organizers matter a great deal, but not mainly through rules about behavior. What works best for them is choosing which problems they can actually check.
The clearest example is ICLR 2026. Program chairs had LLM-detection tools but did not treat them as automatic filters, because the detectors were too unreliable. Detector flags went to human area chairs as one piece of evidence among several. The chairs did desk-reject papers with confirmed fabricated references, because a fake citation is something you can verify (How can conferences detect and handle LLM misuse in peer review?). Compare ICML 2026, where organizers ran a randomized experiment: some reviewers were banned from using LLMs and others were allowed limited use. Scores and decisions barely moved, and many reviewers broke whichever rule they were given (Does banning LLM use in peer review change review outcomes?). Taken together, the two cases suggest a lesson: organizers do better enforcing checkable facts than policing how people work.
The less comfortable finding is that gatekeepers can themselves be the problem. Research on paper mills documents coordinated networks of brokers and editors across countries who pass submissions among themselves. Some editor groups show retraction rates above 50%. When a journal loses its indexing, these networks move to another one (Does scientific fraud operate through organized networks or individual actors?). Here the problem isn't organizers missing fraud. Part of the organizing layer has been captured by it. Another position paper makes a milder version of the same point for AI conferences: review failures are shared among authors, reviewers and venues. It proposes structural fixes, such as letting authors rate a review before they see the decision and giving reviewers badges for thorough work (Can two-stage review and badges fix AI conference peer review?).
Organizers are also being bypassed or gamed from new directions. Eighteen arXiv manuscripts contained hidden instructions telling AI reviewers to rate them favorably (Are hidden AI prompts in preprints a deceptive research practice?). Simply rewriting a paper's text raised AI review scores without improving the science, and AI reviewers agree with each other more than human reviewers do (Can AI systems safely replace human peer reviewers?). A venue that leans on AI reviewing opens a new way for weak papers to get through. Apparent biases also need care before organizers act on them. LLM-assisted reviewers seemed to favor LLM-written papers, but the effect disappeared once paper quality was controlled for. Those reviewers were simply more lenient toward weaker work in general (Do LLM reviewers actually favor LLM-written papers?). And some papers never pass through organizers at all. MIT's case shows an unreviewed preprint shaping public debate long before anyone questioned how reliable it was (Can unreviewed preprints shape scientific debate before peer review?).
One further detail: organizers can also use their venue as an experiment. At Agents4Science, which accepted AI-assisted research, organizers required authors to disclose how much AI was involved. They found that accepted papers had more human guidance than rejected ones (Do accepted papers need more human guidance than rejected ones?). The corpus has little systematic data on how often organizer decisions directly cause bad acceptances. What it does show is organizers working as system designers. Their leverage lies in choosing which checks are verifiable, which incentives to set, and what to measure.
Sources 9 notes
Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.
A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.
Richardson et al. document organized paper mill operations with shared image banks, coordinated editor networks across countries, and strategic journal-hopping when publications lose indexing. Evidence includes 2,213 articles with duplicate images and editor groups exchanging submissions with over 50% retraction rates.
Authors, reviewers, and venues all contribute to peer review failures at major AI conferences. A proposed two-stage system lets authors rate review quality before seeing verdicts, and a badge system rewards reviewer thoroughness, targeting measured biases like rating-length correlation.
Eighteen arXiv manuscripts contained concealed instructions directing AI reviewers to give positive assessments. The practice qualifies as questionable research conduct because concealment plus self-serving design violates ethics regardless of stated intent.
Show all 9 sources
AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.
Across 125,000+ reviews, the apparent favoritism of LLM-assisted reviewers toward LLM papers disappears once paper quality is held constant. LLM papers cluster among weaker submissions, creating a spurious interaction driven by LLM reviewers' general leniency toward lower-quality work.
MIT's case demonstrates that an arXiv preprint shaped AI and science discussions extensively despite never undergoing peer review. When the institution later raised reliability concerns, the damage to discourse had already occurred.
At Agents4Science, organizers observed that accepted papers carried more human input than rejected ones, with humans concentrated in design and hypothesis work while AI gained autonomy in analysis and writing. The pattern emerged from self-reported disclosure tiers across four research stages.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing
- Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
- Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026
- LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on ArXiv
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025