INQUIRING LINE

Citation checkers can wrongly flag real, hard-to-find old sources as fake, and the evidence never measures how often.

What false positive rate do citation verification tools produce on archival works?

This explores how often automated citation checkers wrongly flag real but hard-to-find sources, such as old, out-of-print or non-digitized works, as hallucinated, and what the collection can say about that error rate.


This explores how often citation-checking tools wrongly mark genuine archival sources as fake. The short answer: the collection has no measured false positive rate for archival works, or for any specific category of citation. The closest evidence is GPTZero's scan of 4,841 accepted NeurIPS 2025 papers. It flagged hundreds of possibly hallucinated citations, but every flag still needed a human to confirm it, and the share of flags that held up across the full scan was never published How many accepted conference papers contain hallucinated citations?. That missing number is the key gap. Without it, we can't tell how many flags were fabrications and how many were real references the tool simply couldn't find.

The collection does explain why archival works are where this error is most likely. Most checkers decide whether a reference exists by searching for it. Old material is often poorly digitized, cited in inconsistent formats, or garbled by OCR (the software that turns scanned pages into text). Work on RAG systems built over noisy historical newspapers points to a useful design principle: when the evidence is degraded, the system should refuse to give an answer rather than guess, accepting less coverage in exchange for integrity Can RAG systems refuse to answer without reliable evidence?. Applied to citation checking, that means a tool should report "could not verify" rather than "fabricated". Treating "not found" as "fake" is how archival false positives happen.

A second issue is matching. A real reference may show up with a slightly different title, year or author spelling than the indexed record. A fabricated one may look almost identical to a real paper. Telling these apart is a separate task from finding topically similar documents. Research on two-stage retrieval shows that a dedicated verifier, which compares two texts word by word instead of through compressed summary vectors, can reliably reject near-misses that look structurally similar Can verification separate structural near-misses from topical matches?. A checker that only scores similarity will either accept convincing fakes or reject real citations whose metadata doesn't match exactly.

The errors also run in both directions. Fabrication is now cheap: one demonstration produced 288 finance papers with invented justifications and made-up citations Can AI generate hundreds of fake academic papers automatically?. Generation is outpacing verification across the whole research process Can AI verify research outputs as fast as it generates them?. Meanwhile, readers trust responses with more citations whether or not those citations are relevant Do users trust citations more when there are simply more of them?, and LLM judges give higher scores to answers that include fake references Can LLM judges be tricked without accessing their internals?. So a tool that is too cautious lets fakes through, and one that is too aggressive penalizes scholars who cite obscure archives.

The more useful lesson may concern how results are shown rather than a single error rate. In an 81-person study, readers given no source information couldn't tell truth from fabrication at all. Showing them which claims had been verified restored that ability Can readers tell truth from fabrication without evidence signals?. That suggests citation checkers should report three outcomes (verified, unverifiable and contradicted) rather than a yes/no judgment. Archival works would mostly land in the "unverifiable" group, and that outcome should prompt a human to check rather than serve as an accusation. If you want an actual false positive figure for archival sources, it isn't in this collection yet.


Sources 8 notes

How many accepted conference papers contain hallucinated citations?

GPTZero's citation checker flagged hundreds of potentially hallucinated citations across 4841 accepted NeurIPS 2025 papers. However, flagged citations require human verification to confirm hallucination, and the full verification rate across the full scan remains undisclosed.

Can RAG systems refuse to answer without reliable evidence?

A multilingual RAG system for noisy historical newspapers succeeds by aggressively expanding retrieval while constraining generation to only grounded answers. The grounded-refusal prompt prevents hallucination when OCR errors and language drift degrade source quality, trading coverage for integrity.

Can verification separate structural near-misses from topical matches?

A two-stage pipeline—pooled-cosine recall followed by a small Transformer verifier operating on token-token similarity maps—reliably rejects structural near-misses that MaxSim-style late interaction cannot. The verifier succeeds because it operates on full token interaction patterns rather than compressed vectors.

Can AI generate hundreds of fake academic papers automatically?

A demonstration showed LLMs generating 288 complete finance papers from 96 statistically significant signals, each with invented theoretical justifications and fabricated citations, proving academic HARKing can be automated at scale.

Can AI verify research outputs as fast as it generates them?

AI can produce plausible research outputs faster than it can prove them correct or meaningful, shifting the bottleneck from authorship to verification. Evidence shows 39% of agentic research failures stem from content fabrication and 32% from retrieval failures, not comprehension—and the gap widens precisely where novelty and scientific judgment matter most.

Show all 8 sources
Do users trust citations more when there are simply more of them?

Analysis of 24,000 Search Arena interactions shows irrelevant citations boost user preference (β=0.273) nearly as much as relevant citations (β=0.285), indicating citation count functions as a decoupled trust heuristic.

Can LLM judges be tricked without accessing their internals?

Research shows LLM evaluators systematically score higher when responses include fake references or rich formatting, independent of content quality. These biases are exploitable without model access, undermining AI benchmark credibility.

Can readers tell truth from fabrication without evidence signals?

In an 81-person study, participants given no provenance cues showed no significant truth discernment (p = .43), falling for fluent hallucinations as readily as ground truth. An idealized Provenance Density interface showing verified claims restored a +4.15 point gap (p < .001).

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.