Is the 2024 LLM writing plateau real saturation or measurement artifact?
The adoption curve for LLM-assisted writing flattened in 2024, but the cause remains unclear: either genuine saturation or models becoming too subtle to detect. Resolving this matters for understanding actual usage trends versus measurement limitations.
The excerpt ends its adoption curve on a plateau and names two causes it cannot separate. The abstract puts it as growth that "appears to have stabilized by 2024, reflecting either saturation in LLM adoption or increasing subtlety of more advanced models." The discussion adds domain barriers, such as "costs of adoption, regulatory constraints, concerns over authenticity coupled with advances in users recognizing AI writing," and restates the subtlety explanation, in which AI text becomes "increasingly indistinguishable from human writing, complicating our ability to measure ongoing adoption."
The two readings differ in what a flat line means. If the plateau is saturation, the flattening is the adoption curve, and the same instrument would show it. If the text has become harder to classify, part of the flattening is a property of the measurement. On the excerpt's own logic, a framework that detects LLM-modified text would detect less of it as models improve, so the true share could still be rising while the estimate levels off. The excerpt treats the second reading as a complication for measurement, not as a finding. It describes the framework as validated in earlier work, but it does not say what models that validation covered.
The excerpt positions its framework against commercial detectors and earlier single-domain studies that "relied on black-box commercial AI detectors." A detector-based plateau would face the same subtlety problem and would be harder to audit. Separating the readings needs a measure that does not depend on how detectable the final text is. Can process data distinguish AI delegation from ordinary collaboration? reads contribution timing from process data rather than classifying finished text. In principle that avoids the classification problem, though neither source tests it against this question. Against the How fast did LLM writing adoption actually spread? insight, this question is what the surge note cannot resolve.
The excerpt does not establish which reading holds. It gives no model-capability data over the period and no accuracy figure for the framework on newer text. The implication is that a 2024 plateau in these estimates should be reported as a flattening of the measured share, not as evidence of saturation. Choosing between the readings would need a later window, or an instrument whose accuracy on newer models is known.
Inquiring lines that read this note 2
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How can we detect and account for LLM involvement in academic writing?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
How fast did LLM writing adoption actually spread?
Does LLM-assisted writing use follow a predictable adoption curve across different sectors? Understanding the speed and pattern of adoption helps explain how quickly new AI tools reshape professional communication.
the surge and plateau this question asks about
-
Can process data distinguish AI delegation from ordinary collaboration?
When students or writers use AI tools, their work leaves traces in keystroke logs and editor telemetry. Can these process signatures reliably separate wholesale delegation from permitted collaborative use?
a process-level signal that does not rest on classifying final text
-
Why do newer AI models diverge further from human writing patterns?
As language models improve, they seem to generate text that is measurably less human-like in lexical patterns, yet humans struggle to detect this difference. What drives this divergence, and what does it reveal about how models optimize for quality?
evidence for the harder-to-detect reading: newer LLMs diverge further from human lexical patterns yet human judges find them less detectable
-
Can human judges detect measurable differences in AI text?
Research shows LLM text differs statistically across six lexical dimensions, but human readers—even experts—cannot reliably identify which texts are AI-generated. Why does measurement succeed where human perception fails?
evidence for the too-subtle-to-detect reading: LLM text differs measurably from human text on lexical diversity, though human judges cannot detect the differences
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- The Widespread Adoption of Large Language Model-Assisted Writing Across Society
- Measuring AI "Slop" in Text
- Has the Creativity of Large-Language Models peaked? —an analysis of inter- and intra-LLM variability —
- LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on ArXiv
- Mapping the Increasing Use of LLMs in Scientific Papers
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- Measuring what Matters: Construct Validity in Large Language Model Benchmarks
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
Original note title
the 2024 plateau in LLM-assisted writing may reflect saturation or harder-to-detect text — the measurement cannot tell which