GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
Source: GPTZero · 2026-01-21
ICLR, NeurIPS, ICML, and AAAI are the top machine learning / artificial intelligence conferences in the world, drawing thousands of submissions and participants annually. However, a submission tsunami fueled by generative AI, paper mills, and publication pressure has strained these conferences' review pipelines to the breaking point. Between 2020 and 2025, submissions to NeurIPS increased more than 220% from 9,467 to 21,575. In response, organizers have had to recruit ever greater numbers of reviewers, resulting in issues of oversight, expertise alignment, negligence, and even fraud.
Our purpose in publishing these results is to illuminate a critical vulnerability in the peer review pipeline, not criticize the specific organizers, area chairs, or reviewers who participated in NeurIPS 2025. Over the past several years NeurIPS has changed the review process several times to address problems created by submission volume and generative AI tools. Still, our results reveal the consequences of a system that leaves academic reviewers, editors, and conference organizers outnumbered and outgunned — trying to protect the rigor of peer review against challenges it was never designed to defend against.
These NeurIPS papers have already been accepted, presented live, and effectively published. Since NeurIPS 2025 had an acceptance rate for main track papers of 24.52%, each of these papers beat out 15,000 other papers despite containing one or more hallucinations. This is concerning, given that the NeurIPS LLM policy considers hallucinated citations to be grounds for a paper's rejection or revocation, similar to ICLR.
We've scanned each paper for both hallucinated citations (Sources) and AI-generated text (AI). An "*" next to the scan indicates the paper is likely a mix of AI and human text, while "**" indicates the paper is likely AI-generated.
Given the high stakes for both authors and publishers, GPTZero's Hallucination Check is engineered to be accurate, transparent, and cautious. It uses our AI agent, trained in-house, to flag any citations in a document that can’t be found online. These flagged citations are not automatically hallucinations — many archival documents or unpublished works can’t be matched to an online source — but they indicate which sources require further human scrutiny. As always, we recommend that a human confirm that flagged citation is an AI-generated fake instead of the result of a more conventional error.
We define a vibe citation as a citation that likely resulted from the use of generative AI. Vibe citing results in errors common to LLM generations, but rare in human-written text, such as:
Fabricating the author(s), title, URL/DOI, and/or container (ex.
Modifying the author(s) or title of a source by extrapolating a first name from an initial, dropping and/or adding authors, or paraphrasing the title.
Like GPTZero’s AI Detector, Hallucination Check has an extremely low false negative rate, so we catch 99 out of 100 flawed citations. Because our tool will flag any citation that can't be verified online, the false positive rate is higher.
GPTZero's analysis of 4841 of the 5290 papers accepted by NeurIPS 2025 indicates noticeable traces of AI authorship and hundreds of vibe citations. As always, each of the hallucinations presented here has been verified by a human expert.
Hallucination Check is the only tool of its kind, and provides an essential service at multiple points in the peer review pipeline. First, it allows authors to check their manuscripts for citation errors — including common issues that can occur without LLM involvement like dead links or partial titles. Second, it greatly reduces the time and labor necessary for reviewers to check a submission's sources and identify possible vibe citing. Third, using Hallucination Check in combination with GPTZero's AI Detector allows editors and conference chairs to check for AI-generated text and suspicious citations at the same time, leading to faster and more accurate editorial decisions.
After releasing our ICLR paper investigation we are now coordinating with the ICLR team to review future paper submissions. As always, our goal is to make the peer review process faster, fairer, and more transparent for everyone involved. Try GPTZero's Hallucination check for yourself, or reach out to GPTZero's team.
Lines of inquiry this paper opens 23
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can AI systems perform peer review as effectively as humans? How do hallucinated citations emerge in AI scholarly output?- What false positive rate do citation verification tools produce on archival works?
- How much undetected fraud exists beyond current retraction statistics?
- How do courts distinguish between AI hallucinations and ordinary typographical errors?
- How often do fabricated sources in AI output escape citation checking?
- How many NHANES studies apply false discovery correction before publication?
- Do depression associations in NHANES papers survive correction for multiple comparisons?
- Do solo lawyers face different citation hallucination risks than large firms?
- What percentage of AI hallucination cases result in actual court sanctions?
- Do solo practitioners face different sanction outcomes than large law firms?
- What standard of intent or bad faith triggers Rule 11 sanctions for citations?
- How much do existing legal AI tools actually hallucinate in practice?
- How do different legal AI tools compare in accuracy across case eras?
- What happens when lawyers rely on AI citations that turn out false?
- Do certain news topics trigger more hallucinations than others?
- Does retrieval augmented generation actually eliminate hallucinations in any domain?
- Can architectural changes reduce hallucination without external retrieval or verification?
- Why do hallucination rates differ between vendor AI products and student-used models?
- Do legal AI tools marketed as hallucination-free actually hallucinate?