Where is most LLM-generated content actually appearing in computer science?
Review papers show higher LLM-generated shares than other papers, but the raw volume tells a different story. This note explores which paper types actually contain most generated content by count.
The authors test the rationale for arXiv's October 31, 2025 policy, which bars unpublished review, survey and position papers from its CS servers. arXiv cited rising LLM-generated content in reviews but gave no data. Here, review papers carry a higher share of LLM-generated content: 21.4% against 14.0% for non-review papers, as adjusted cohort-level estimates over three years. Counted by paper, the volume sits mostly outside reviews. For 2025 the authors estimate 4,783 LLM-generated review papers against 26,801 non-review papers, almost six times as many, out of about 12K review papers and 141K non-review papers. Using the commercial detector Pangram, they put 2025 shares at 43.3% for reviews and 23.3% for non-review papers.
The estimates are population-level by design; the authors do not claim to label individual papers. A classifier first separates review from regular papers, scoring 92.0% F1 on 200 hand-annotated CS papers, half selected for review keywords. Two detectors then estimate the LLM-generated fraction of each group: the Alpha estimator (Liang et al., 2024a), which relies on adjective occurrence and needed a type-specific Rogan-Gladen correction for different false-positive rates, and Pangram, which the authors report produced no false positives. The main pipeline uses titles and abstracts, the same data arXiv moderators saw. The authors add that "review" is not a neutral category, since the annotation follows a "typical" CS view rather than each subfield's norms.
The temporal result sharpens its contrast with How fast did LLM writing adoption actually spread?. That note reports a surge followed by a plateau across four public-facing domains; this excerpt describes "an accelerating trend since 2023" across arXiv's CS, Physics, Mathematics and Statistics. The corpora and estimators differ, so neither settles the other, but they disagree on whether adoption has leveled off. On policy, Does banning LLM use in peer review change review outcomes? measures what a restriction does to review outcomes, whereas this paper measures where generated text sits and finds reviews a minority of it. On detection, Why did AI article share stop growing after 2025? raises the open question of whether a detector-based share reflects use or what the detector can see; this paper addresses that with a correction rather than resolving it.
The excerpt does not establish that generated papers are weaker; it measures how much generated text appears, not its quality. It contains no moderation data, so the claim that the ban "makes little sense" for saving moderator time is an inference from paper counts. The counts also inherit detector uncertainty: the Alpha estimate may shift with the generating model, the estimation data or human post-editing, and the classifier's borderline cases, such as tutorials and hybrid papers, stay ambiguous. The supported implication is narrower than a verdict on the policy: a review-only restriction would miss most of the generated volume, at roughly the magnitude these estimates give. Whether it should be replaced depends on moderation evidence this excerpt does not include.
Inquiring lines that read this note 2
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How can we detect and account for LLM involvement in academic writing? How can AI systems reliably guide voters without introducing political bias?Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
How fast did LLM writing adoption actually spread?
Does LLM-assisted writing use follow a predictable adoption curve across different sectors? Understanding the speed and pattern of adoption helps explain how quickly new AI tools reshape professional communication.
its plateau finding contrasts with this excerpt's accelerating trend in academic CS and sciences.
-
Does banning LLM use in peer review change review outcomes?
Can policies restricting or allowing AI tools shift how reviewers score papers and make decisions? This matters because review quality and fairness depend on consistent standards.
also studies a restriction on LLM use, but measures review outcomes where this paper measures prevalence.
-
Why did AI article share stop growing after 2025?
Graphite reports a plateau in AI-generated articles at 50% but cannot distinguish whether search engines penalize AI content, detection tools are missing more AI text, or both. The excerpt leaves this critical cause untested.
shares the detector-dependence problem; this paper addresses it with a type-specific correction.
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on ArXiv
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- Pangram Predicts 21% of ICLR Reviews are AI-Generated
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- Has the Creativity of Large-Language Models peaked? —an analysis of inter- and intra-LLM variability —
- Can large language models provide useful feedback on research papers? A large-scale empirical analysis
- Mapping the Increasing Use of LLMs in Scientific Papers
- LLM-REVal: Can We Trust LLM Reviewers Yet?
Original note title
review papers carry a higher LLM-generated share than non-review papers — but non-review papers yield almost six times as many LLM-generated papers in CS