Most studies using a big US health survey skip a standard false-alarm check, and in one test only 13 of 28 findings held up.
How many NHANES studies apply false discovery correction before publication?
This explores how often papers built on NHANES (the large US national health survey dataset) correct for multiple comparisons, meaning they check whether 'significant' findings would survive testing many hypotheses at once, before they get published.
This explores how often NHANES-based studies correct for multiple comparisons before publication. The collection doesn't give an exact count. It has one systematic review on the question, which says most of these papers skip the correction. That review Why did single-factor NHANES studies explode after 2021? found 341 single-factor NHANES papers over a decade. Output averaged about four a year until 2021, then jumped to 190 in 2024. Most of these papers left out false discovery correction and didn't model more than one factor at a time. The clearest test came from the 28 depression studies: once the reviewers applied the correction themselves, only 13 of the 28 reported associations held up. So the honest answer is 'a minority, and the uncorrected results are fragile.' For a precise percentage, you'd need to go to the review itself.
The number matters less than the pattern behind it. A large public dataset, one exposure tested against one outcome, and no correction for the many other comparisons that could have been run: that combination produces results that look publishable and often don't replicate. When output jumps from four papers a year to almost two hundred, it suggests the cost of writing such a paper fell faster than the cost of checking it. The review documents the jump and the design gaps. It doesn't establish what caused the jump.
The broader collection keeps coming back to this gap between producing research and verifying it. One approach to AI-assisted paper writing Can separating judgment from verification improve research paper reliability? requires authors to say what evidence they expect before they see any results. That is close to a structural fix for the problem the NHANES review describes, because deciding the analysis in advance takes away the freedom to search for whichever comparison comes out significant. A critique of ad hoc prompt engineering Does iterative prompt engineering undermine scientific validity? makes a similar point from another field: if you keep revising your method until the output looks right, you've built a self-fulfilling loop, and the remedy is to fix your criteria ahead of time.
On the review side, the hopeful finding is that automated checking can catch what human reviewers miss. An agentic reviewer that spends extra computation checking proofs and experiments line by line Can inference scaling help reviewers catch errors humans miss? found flaws in papers that had already been accepted at top venues. Missing multiple-comparison correction is exactly the kind of mechanical, checkable gap such tools could flag. The less hopeful findings show how much already gets through. Hundreds of possibly fabricated citations turned up in accepted NeurIPS papers How many accepted conference papers contain hallucinated citations?, and some authors have hidden instructions in their papers telling AI reviewers to be generous Are hidden AI prompts in preprints a deceptive research practice?.
The NHANES case doesn't really involve AI. It's a story about publishing volume outgrowing statistical discipline. But it's the clearest example in the collection of something easy to forget: a basic, decades-old statistical safeguard can drop out of practice quietly, and you only notice when someone reruns the analysis and half the findings disappear.
Sources 6 notes
A systematic review found 341 NHANES papers over a decade, with volume jumping sharply after 2021—most omitting false discovery correction and multifactorial modeling. Among 28 depression studies, only 13 of 28 associations survived multiple-comparison correction.
Spark-to-Paper architects paper generation as composable skills that isolate model judgment from executable, verifiable operations and require evidence specification before results are observed, reducing dependence on model correctness for consistency.
Iterative prompt revision by single researchers introduces individual bias, shifts evaluation criteria to match LLM capabilities rather than task requirements, and creates self-fulfilling feedback loops. A validated pipeline with inter-coder reliability and pre-specified criteria is required instead.
PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.
GPTZero's citation checker flagged hundreds of potentially hallucinated citations across 4841 accepted NeurIPS 2025 papers. However, flagged citations require human verification to confirm hallucination, and the full verification rate across the full scan remains undisclosed.
Show all 6 sources
Eighteen arXiv manuscripts contained concealed instructions directing AI reviewers to give positive assessments. The practice qualifies as questionable research conduct because concealment plus self-serving design violates ethics regardless of stated intent.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Towards Automating Scientific Review with Google's Paper Assistant Tool
- Stop Automating Peer Review Without Rigorous Evaluation
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
- Hidden Prompts in Manuscripts Exploit AI-Assisted Peer Review
- From Prompt Engineering to Prompt Science With Human in the Loop
- Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill
- Scholars sneaking phrases into papers to fool AI reviewers