Why did single-factor NHANES studies explode after 2021?
NHANES papers proposing one-predictor associations surged from 4 per year to 190 in 2024, raising questions about what enabled the volume spike and whether design shortcuts became systematically common.
Suchak et al. report a systematic literature search that found 341 NHANES-derived papers over the past decade, each proposing an association between one predictor and one health condition. The volume is the headline: "an average of 4 single-factor manuscripts ... per year between 2014 and 2021," then "190 in 2024 up to 9 October." The authorship shift is as sharp: 2 of 25 manuscripts from 2014 to 2020 had a primary author affiliated in China, against 292 of 316 from 2021 to 2024. The depression studies are the authors' test of the statistics. Of 28 depression papers, none applied false discovery correction. When the authors applied Benjamini-Yekutieli correction across the 28 associations, "less than half (13) remained statistically significant."
The mechanism the authors give is about how cheap the work has become. NHANES is an "AI-ready dataset" that can be pulled via API straight into R or Python, so the number of hypotheses is "constrained only by computational access." The authors argue that single-factor AI-supported analysis "removes context from research, fails to capture interactions, avoids false discovery correction, and is an approach that can easily be adopted by paper mills." Two further moves multiply output: reversing predictor and outcome, which they call the most extreme case, and selective data use, such as limited date ranges or cohort subsets, which they read as "suggestive of data dredging, and post-hoc hypothesis formation." The statistics are not new. The claim is that an AI-supported pipeline makes a known practice cheap enough to run at scale.
Against the nearest notes, this excerpt documents the failure that Can separating judgment from verification improve research paper reliability? is designed to prevent. Spark-to-Paper keeps checkable operations apart from model judgment and sets required evidence before results are seen. The depression studies skipped exactly the kind of deterministic step, a multiple-comparison correction, that such a pipeline could enforce. This excerpt is an audit of published output; Spark-to-Paper is a builder's design. The two read as complementary. The surge also sharpens How fast did LLM writing adoption actually spread?. That note measures a rise in public-facing text across four domains; this one shows a rise concentrated in one formulaic genre, with defects in the design rather than the sentences. That matters for Can people reliably spot content made by AI?: a reader relying on prose cues would miss these papers, because the authors' signal is structural (single predictors, no correction, shifting windows).
What the excerpt does not establish is the link to AI itself. The authors measure volume, affiliation and analytic design. They do not detect AI use in any of the 341 papers. "AI-assisted productivity" is an inference from the timing of the rise, and the paper-mill connection is framed as a risk and a case study, not a finding. The China-affiliation shift is a bibliometric fact that says nothing about who produced any given paper. The 13-of-28 result is a reanalysis that treats the 28 depression studies as one family of hypotheses, and it depends on that choice. The excerpt also ends before the authors' best-practice recommendations, so those remedies are not assessed here. The defensible reading is narrower than the headline. The rise in single-factor NHANES papers and their missing corrections are documented. That AI tools caused the rise is a hypothesis the excerpt motivates but does not test.
Inquiring lines that read this note 3
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What gaps exist between benchmark performance and real deployment outcomes? What human oversight must AI research systems have?Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can separating judgment from verification improve research paper reliability?
Explores whether dividing model-based decisions from deterministic checks and fixing evidence requirements before observing results could bound errors in automated paper generation and make AI-assisted research more trustworthy.
contrast: a builder's design for enforced checks; this excerpt audits what happens when such checks are absent
-
How fast did LLM writing adoption actually spread?
Does LLM-assisted writing use follow a predictable adoption curve across different sectors? Understanding the speed and pattern of adoption helps explain how quickly new AI tools reshape professional communication.
extends: a surge measured in general public-facing text, here concentrated in one formulaic research genre
-
Can people reliably spot content made by AI?
This systematic review of 30 studies asks whether human judgment can distinguish AI-generated text, images, and voice from human-created content, and whether detection accuracy has improved as AI becomes more realistic.
contrast: these papers are flagged by analytic structure, not by prose cues a reader would catch
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Explosion of formulaic research articles, including inappropriate study designs and false discoveries, based on the NHANES US national health database
- Accelerating science with human-aware artificial intelligence
- Is it Cake or is it AI? A Systematic Review of Human Uncertainty in Distinguishing Generative Artificial Intelligence Content
- GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
- Causal Claims in Economics
- Most peer reviewers now use AI, and publishing policy must keep pace
- How to Find Fantastic AI Papers: Self-Rankings as a Powerful Predictor of Scientific Impact Beyond Peer Review
- Adoption and Impact of Command-Line AI Coding Agents: A Study of Microsoft's Early 2026 Rollout of Claude Code and GitHub Copilot CLI
Original note title
NHANES single-factor papers averaged four a year to 2021 and hit 190 in 2024 — formulaic studies skipping multifactorial models and false discovery correction