SYNTHESIS NOTE
Topics›Domain Specialization›this note

How much GPT-written scholarship reaches Google Scholar undetected?

Haider et al. searched Google Scholar for telltale ChatGPT phrases to estimate how many papers contain undisclosed AI authorship, especially in policy-relevant fields. Understanding prevalence matters because lay readers—politicians, patients, students—may treat these papers as credible research.

Synthesis note · 2026-10-06 · sourced from Domain Specialization

Haider et al. argue that undisclosed GPT-written scientific papers are already inside Google Scholar, listed alongside reputable research. They retrieved papers containing at least one of two phrases that conversational agents commonly return, then coded the sample and ran descriptive statistics. The excerpt gives two figures for the same group: "roughly two-thirds" were produced "at least in part, through undisclosed, potentially deceptive use of GPT," and around 62% did not declare GPT use. Of these papers, 57% concerned policy-relevant subjects: computing (23%), environment (19.5%) and health (14.5%). Most sat in non-indexed journals and working papers, though some appeared in mainstream journals and conference proceedings, and most existed in several copies across domains.

The mechanism the authors give combines leakage with retrieval. Since the public release of ChatGPT in 2022, and given how Google Scholar lists everything together, lay readers such as "media, politicians, patients, students" are more likely to meet these papers, and mirror sites and shadow libraries keep copies circulating after any retraction. The authors also report that joint reading found GPT-produced text across most sections of the papers, leading them to write that "almost all articles in our sample of questionable articles likely contained traces of GPT-fabricated text everywhere." If that holds, the telltale phrases set a floor on detection rather than a measure of extent. They call the strategic version of this risk "evidence hacking," the "strategic and coordinated malicious manipulation of society's evidence base," and tie its greatest danger to politically divisive domains.

Against the nearest notes, this is a different route to the evidence. How much machine-generated text actually appears on Reddit? estimates prevalence with a zero-shot detector over posts; Haider et al. instead start from phrases the generator itself emits and then check copies with ordinary web search. That method finds the unedited output and says nothing about polished output, which the sample cannot reach by construction. It also meets Does polished writing actually signal better quality work? head on: the authors' worry is papers that look "convincingly scientific-looking," and an index that ranks by relevance rather than provenance rewards exactly that look. The contrast with Do AI slop accusations actually detect AI text? is sharp. There, prose features do not predict who gets accused; here, a surface phrase does track GPT output, where the generator's own wording survived. Surface signals work for finding unedited text and are weak for the rest.

The excerpt does not establish how large the problem is. It gives no sample size, and the 62% and 57% are shares of a phrase-matched sample, not of Google Scholar's index or of all GPT-assisted scholarship; a paper that scrubbed the phrases would never enter the sample. The copy counts, the claim that retraction can backfire, and the link to mistrust among people who "do their own research" are argued from prior work and the mirror-site mechanism, not measured here. The data are deposited in the Harvard Dataverse, so the figures can be checked, but the excerpt cannot settle the scale of influence. What the evidence supports at its strength is narrower than the authors' framing: indexed scholarly search does not separate provenance from quality, so a reader cannot rely on the index to flag a fabricated paper, and disclosure at the source is what remains.

Inquiring lines that read this note 2

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do hallucinated citations emerge in AI scholarly output? Can AI systems perform peer review as effectively as humans?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 98 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

Haider et al. find undisclosed GPT-written papers reach Google Scholar, 57 percent on policy topics — the authors call the risk evidence hacking