SYNTHESIS NOTE
Topics›Co Writing Collaboration›this note

Can AI generate hundreds of fake academic papers automatically?

Explores whether language models can industrialize academic fraud by retroactively constructing theoretical justifications for data-mined patterns, complete with fabricated citations and creative signal names.

Synthesis note · 2026-03-27 · sourced from Co Writing Collaboration

A demonstration paper applied LLMs to generate three distinct complete versions of academic papers for each of 96 stock return predictor signals. Each version included "creative names for the signals, custom introductions providing different theoretical justifications for the observed predictability patterns, and citations to existing (and, on occasion, imagined) literature." This is HARKing (Hypothesizing After Results are Known) industrialized.

The process: mine 30,000+ potential predictor signals from accounting data, apply rigorous statistical filtering to find 96 that pass, then use LLMs to retroactively construct theoretical justifications for why those signals should predict returns. The AI generates the narrative that makes the data mining look like hypothesis-driven research.

This is the academic equivalent of the false punditry described in the social media context — style substituting for thought at industrial scale. Since Does polished AI output trick audiences into trusting it?, the generated papers exploit the same heuristic: professional-looking output implies expert-quality thinking. And since Should we call LLM errors hallucinations or fabrications?, the process that generates valid theoretical justifications is identical to the process that generates fabricated ones.

Inquiring lines that read this note 65

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do hallucinated citations emerge in AI scholarly output? Can AI systems perform peer review as effectively as humans? Can humans reliably detect and resist AI-generated misinformation? Why does polished AI output gain credibility despite fundamental verifiability problems? Can artificial systems establish authority in domains requiring expert judgment? How do writers navigate authorship and delegation with AI? What human oversight must AI research systems have? How can evaluations be made robust against model reward hacking? Can AI systems evade safety evaluations through reasoning manipulation? What are the real-world consequences of AI citation hallucinations? Why does AI verification capability persistently exceed generation capability? Can readers reliably distinguish AI-written text from human writing? How reliably can humans and AI detectors identify machine-generated text? How does AI-generated content create social proof without authentic interaction? How do AI hiring systems affect authenticity, fairness, and candidate preferences? How do educators verify student capability when AI can produce indistinguishable work? How can we detect and account for LLM involvement in academic writing? Does AI assistance erode cognitive skills while inflating perceived competence?

Related concepts in this collection 2

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
17 direct connections · 161 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

AI can industrialize hypothesis-after-results-known by auto-generating hundreds of complete academic papers with creative names and citations to imagined literature