How much GPT-written scholarship reaches Google Scholar undetected?
Haider et al. searched Google Scholar for telltale ChatGPT phrases to estimate how many papers contain undisclosed AI authorship, especially in policy-relevant fields. Understanding prevalence matters because lay readers—politicians, patients, students—may treat these papers as credible research.
Haider et al. argue that undisclosed GPT-written scientific papers are already inside Google Scholar, listed alongside reputable research. They retrieved papers containing at least one of two phrases that conversational agents commonly return, then coded the sample and ran descriptive statistics. The excerpt gives two figures for the same group: "roughly two-thirds" were produced "at least in part, through undisclosed, potentially deceptive use of GPT," and around 62% did not declare GPT use. Of these papers, 57% concerned policy-relevant subjects: computing (23%), environment (19.5%) and health (14.5%). Most sat in non-indexed journals and working papers, though some appeared in mainstream journals and conference proceedings, and most existed in several copies across domains.
The mechanism the authors give combines leakage with retrieval. Since the public release of ChatGPT in 2022, and given how Google Scholar lists everything together, lay readers such as "media, politicians, patients, students" are more likely to meet these papers, and mirror sites and shadow libraries keep copies circulating after any retraction. The authors also report that joint reading found GPT-produced text across most sections of the papers, leading them to write that "almost all articles in our sample of questionable articles likely contained traces of GPT-fabricated text everywhere." If that holds, the telltale phrases set a floor on detection rather than a measure of extent. They call the strategic version of this risk "evidence hacking," the "strategic and coordinated malicious manipulation of society's evidence base," and tie its greatest danger to politically divisive domains.
Against the nearest notes, this is a different route to the evidence. How much machine-generated text actually appears on Reddit? estimates prevalence with a zero-shot detector over posts; Haider et al. instead start from phrases the generator itself emits and then check copies with ordinary web search. That method finds the unedited output and says nothing about polished output, which the sample cannot reach by construction. It also meets Does polished writing actually signal better quality work? head on: the authors' worry is papers that look "convincingly scientific-looking," and an index that ranks by relevance rather than provenance rewards exactly that look. The contrast with Do AI slop accusations actually detect AI text? is sharp. There, prose features do not predict who gets accused; here, a surface phrase does track GPT output, where the generator's own wording survived. Surface signals work for finding unedited text and are weak for the rest.
The excerpt does not establish how large the problem is. It gives no sample size, and the 62% and 57% are shares of a phrase-matched sample, not of Google Scholar's index or of all GPT-assisted scholarship; a paper that scrubbed the phrases would never enter the sample. The copy counts, the claim that retraction can backfire, and the link to mistrust among people who "do their own research" are argued from prior work and the mirror-site mechanism, not measured here. The data are deposited in the Harvard Dataverse, so the figures can be checked, but the excerpt cannot settle the scale of influence. What the evidence supports at its strength is narrower than the authors' framing: indexed scholarly search does not separate provenance from quality, so a reader cannot rely on the index to flag a fabricated paper, and disclosure at the source is what remains.
Inquiring lines that read this note 2
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do hallucinated citations emerge in AI scholarly output? Can AI systems perform peer review as effectively as humans?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does polished writing actually signal better quality work?
When evaluators judge applications and manuscripts, does rhetorical sophistication predict merit, or does it distract from verifiable evidence of competence and rigor?
polish as a quality signal is what fabricated papers exploit at scale in an index that ranks by relevance.
-
How much machine-generated text actually appears on Reddit?
Researchers ran a detector across millions of Reddit posts and comments to measure how prevalent AI-written content is on the platform. Understanding this prevalence matters for assessing Reddit's authenticity and the scale of AI adoption in online communities.
contrasts a classifier-based prevalence estimate with this phrase-matched sample, which cannot give a base rate.
-
Do AI slop accusations actually detect AI text?
When online communities label comments as AI-generated slop, are they identifying genuine machine writing or enforcing social boundaries? This asks whether the accusation register tracks real detection or functions as gatekeeping.
contrast: prose features fail to predict accusation, while a surface phrase tracks GPT output here.
-
Can people reliably spot content made by AI?
This systematic review of 30 studies asks whether human judgment can distinguish AI-generated text, images, and voice from human-created content, and whether detection accuracy has improved as AI becomes more realistic.
readers near chance argue for checks on the search side, which is where this paper finds the papers.
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- GPT-fabricated scientific papers on Google Scholar: Key features, spread, and implications for preempting evidence manipulation
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- Hidden Prompts in Manuscripts Exploit AI-Assisted Peer Review
- Mapping the Increasing Use of LLMs in Scientific Papers
- Are We in the AI-Generated Text World Already? Quantifying and Monitoring AIGT on Social Media
- Machines in the Crowd? Measuring the Footprint of Machine-Generated Text on Reddit
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- ChatGPT: deconstructing the debate and moving it forward
Original note title
Haider et al. find undisclosed GPT-written papers reach Google Scholar, 57 percent on policy topics — the authors call the risk evidence hacking