INQUIRING LINE

After AI writing surged, its share of public text leveled off. Did people stop using it, or did detectors stop catching it?

Why did LLM-assisted writing plateau instead of continuing to grow?

This explores why measured use of AI in public writing (complaints, press releases, job posts, UN documents) rose quickly after ChatGPT's launch and then leveled off, and whether that leveling-off is real or a side effect of how it was measured.


This explores why the share of AI-assisted text in public writing stopped climbing after its early surge, and whether that flat line means what it appears to mean. The pattern itself is well documented. Across four very different kinds of writing (consumer complaints, corporate press releases, job postings, and UN documents), LLM-assisted text rose sharply in the months after ChatGPT appeared, then settled at roughly 10–24% by 2024. Smaller and younger organizations moved first How fast did LLM writing adoption actually spread?. The obvious reading is saturation: everyone who was going to use it already did.

The corpus doesn't let you settle on that reading. The plateau fits two quite different explanations. People may really have stopped adopting, or newer models may write text that the detection method can no longer spot. The study didn't have the model-capability data or accuracy figures on newer text needed to tell these apart Is the 2024 LLM writing plateau real saturation or measurement artifact?. So the honest answer to "why did it plateau?" is that nobody knows yet whether it did.

Other research in the collection makes the second explanation plausible. Studies of AI editing documents find that weaker models fail visibly, by deleting content. Frontier models fail quietly instead: they corrupt meaning while keeping the surface looking intact Does model capability change how documents degrade?. If stronger models leave fewer visible traces when they go wrong, they probably leave fewer stylistic fingerprints when they go right, and fingerprints are what word-frequency detectors rely on. A similar lesson comes from a different area: the sudden 'emergent abilities' of large models largely disappear when you switch to a smoother metric Are LLM emergent abilities real or measurement artifacts?. In both cases, the shape of the curve can come from the measuring tool as much as from the thing being measured.

There's also a reason to doubt that real use leveled off. AI-written text gets rewarded when another AI is doing the reading. LLMs evaluating resumes prefer their own rewrites over equivalent human versions, at rates up to 98%, and the bias grows in larger models Do language models favor resumes they rewrote themselves?. In simulated hiring pipelines, applicants who used the same model as the screener were shortlisted 23–60% more often Do LLM evaluators favor resumes written by their own model?. Where AI screens the writing, people have a growing reason to use it, which makes a natural stopping point less likely.

The takeaway you may not have expected: a flat line in 'how much text is AI-written' is partly a measure of how detectable AI writing is. As models get better at sounding human, adoption curves built on detection will tend to flatten whether or not use does. The corpus lacks the direct evidence, such as detector accuracy on 2024-era models, that would settle the question.


Sources 6 notes

How fast did LLM writing adoption actually spread?

Across consumer complaints, press releases, job postings, and UN documents, LLM-assisted text rose sharply in the months after ChatGPT's launch and stabilized at roughly 10–24% by 2024. Smaller and younger organizations adopted faster than larger ones.

Is the 2024 LLM writing plateau real saturation or measurement artifact?

A plateau in LLM-assisted writing estimates could reflect either genuine adoption saturation or improved model subtlety that evades detection. The framework lacks model-capability data and accuracy figures on newer text to separate these readings.

Does model capability change how documents degrade?

DELEGATE-52 shows weaker LLMs degrade documents through visible deletion, while frontier models degrade through subtle corruption that preserves surface integrity. This shift makes frontier failures harder to detect and potentially more dangerous at workflow scale.

Are LLM emergent abilities real or measurement artifacts?

Sharp, unpredictable capability transitions vanish when using continuous metrics instead of discontinuous ones. The same model outputs show smooth predictable improvement with scale, suggesting emergence is a measurement choice rather than a real behavioral change.

Do language models favor resumes they rewrote themselves?

Across a controlled experiment on 2,245 resumes, eight of nine LLMs preferred their own rewrites over matched human versions when evaluating candidates, with preference rates ranging from 26% to 98%. The bias strengthened in larger models and emerged from stylistic alignment rather than content quality differences.

Show all 6 sources
Do LLM evaluators favor resumes written by their own model?

Simulations across 24 occupations show applicants using the evaluating LLM are significantly more likely to advance past resume screening than equally qualified human-written applicants, with the largest gaps in business fields like sales and accounting.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.