Two years after ChatGPT, detectors rated about half of web articles mostly AI-written, then growth stalled, perhaps because AI pages rank poorly in search.
Did LLM-generated content surge after ChatGPT or has it plateaued?
This explores whether the share of AI-written text online kept climbing after ChatGPT's release in late 2022 or has levelled off, and what the corpus can measure about that.
This explores whether AI-written text kept spreading after ChatGPT launched or has stopped growing. The corpus has one direct measurement for the open web, and it shows both: a fast rise, then a plateau. Graphite ran three AI detectors over 55,400 web articles and averaged their results. Articles that were mostly AI-written went from almost none to about 36% within a year of ChatGPT and reached 48% by the two-year mark. Since early 2025 the share has held at roughly half Has AI-generated content stopped growing on the web?. The reason it levelled off is the interesting part. The suggested explanation is that AI articles rank poorly in search, so producing more of them earns little. If that's right, the ceiling comes from economics, not from people losing interest in AI writing. Keep in mind this is one study built on detectors, and detectors are imperfect, so treat the 50% figure as an estimate rather than a census.
Research papers tell a different story than the web. In computer science papers, the AI-generated share is measured in proportions of text, not whole articles. Review papers show more of it (about 21%) than ordinary research papers (about 14%). But ordinary papers are far more numerous, so they account for nearly six times as many AI-generated papers overall: about 26,800 against 4,800 in a 2025 estimate Where is most LLM-generated content actually appearing in computer science?. So asking 'where did the surge land?' matters as much as asking 'how big was it?' The heaviest users by percentage are not where most of the volume sits.
The spread also reaches peer review, and here the obvious story turns out to be wrong. Across more than 125,000 reviews, reviewers who used AI seemed to favor AI-written papers. The effect disappears once paper quality is taken into account. AI-assisted papers tend to be weaker, and AI-assisted reviewers are generally easier on weaker work Do LLM reviewers actually favor LLM-written papers?. The lesson carries over to any surge statistic: raw shares can hide what is actually driving them.
What you might not have expected to care about is that a 50% plateau is not a steady state for the models themselves. One argument in the corpus holds that when LLMs improve, they are mostly picking up shifts in human-written text, not reasoning better Do LLM improvements reflect reasoning gains or corpus shifts?. Another treats model output as a draw from the model's learned assumptions, not as an observation of the world Should we treat LLM outputs as real empirical data?. Put these together and a web that is half AI-written becomes a feedback problem: future models trained on it partly learn from earlier models' outputs. Generation also tends to produce smooth, familiar claims rather than new viewpoints Does LLM generation explore competing claims while producing text?, so a plateau in quantity could still mean a slow decline in variety. The corpus doesn't measure that decline directly. That is the open question these notes point toward.
Sources 6 notes
Graphite's analysis of 55.4k web articles found AI-primary content rose to 35.9% within 12 months of ChatGPT, reached 48% by month 24, then stabilized near 50% from Q1 2025 onward. The plateau may reflect poor search performance of AI articles.
Review papers contained 21.4% LLM-generated content versus 14.0% for non-review papers in adjusted estimates, yet non-review papers yielded almost six times more generated papers by volume. A 2025 estimate found 4,783 generated review papers against 26,801 generated non-review papers.
Across 125,000+ reviews, the apparent favoritism of LLM-assisted reviewers toward LLM papers disappears once paper quality is held constant. LLM papers cluster among weaker submissions, creating a spurious interaction driven by LLM reviewers' general leniency toward lower-quality work.
The research argues that when LLMs succeed on tasks they once failed, they are mimicking shifts in human-written discourse rather than developing new reasoning abilities. The intelligence resides in humans and the corpus, not in the models themselves.
Foundation Priors framework shows that LLM-generated text reflects the model's learned patterns and user's prompt choices, not ground truth. Such outputs should only influence inference through explicitly parameterized trust weights, not be treated as equivalent to real evidence.
Show all 6 sources
Token prediction trains models to continue toward the training distribution, not to explore logically related counterpositions. This smoothness in process produces smooth claims that multiply without generating new perspectives.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- Measuring Faithfulness in Chain-of-Thought Reasoning
- Argument Collapse: LLMs Flatten Long-Form Public Debate
- LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on ArXiv
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- Has the Creativity of Large-Language Models peaked? —an analysis of inter- and intra-LLM variability —
- AI Now Writes as Many Online Articles as Humans