SYNTHESIS NOTE
Topics›Domain Specialization›this note

Does LLM writing assistance change how scientists publish?

When scientists adopt LLMs to draft manuscripts, do they produce more papers, and does writing quality still signal merit? This matters because it affects how we evaluate scientific work.

Synthesis note · 2026-10-06 · sourced from Domain Specialization

Across arXiv, bioRxiv and SSRN, about 2.1M preprints from January 2018 to June 2024, the authors find that scientists who adopt LLMs to draft manuscripts show "a large increase in paper production, ranging from 23.7-89.3% depending on scientific field and author background." The same abstract reports that LLM use "has reversed the relationship between writing complexity and paper quality," producing "an influx of manuscripts that are linguistically complex but substantively underwhelming." It also reports that adopters access and cite more diverse prior work, including books and younger, less-cited documents. These are the authors' own measurements. The LLM-use classification is their text-based algorithm, not a commercial detector, and the production range is their estimate.

The mechanism the authors give is the cost of writing. A distinctive paper requires "compelling written arguments" and links to prior literature, and these tasks are "time consuming, particularly for researchers communicating in a non-native language." LLMs, they argue, "should asymmetrically reduce the cost of writing across scientists' linguistic backgrounds," which is why they expect productivity gains to concentrate among researchers facing higher writing costs. On quality, the argument rests on a heuristic: clear but complex prose has been read as a sign of merit because producing it used to take effort. Once that effort falls, polished prose stops being a reliable "signal of an author's command of a topic," and for LLM-assisted manuscripts the complexity-merit correlation "not only disappears; it inverts." The detector compares the token distribution of pre-2023 abstracts with that of GPT-3.5-turbo rewrites of them, and uses the comparison to identify probable LLM-assisted abstracts.

Against the nearest notes, this excerpt extends two existing claims to scientific preprints. How fast did LLM writing adoption actually spread? uses a similar detection-based measurement on consumer complaints, press releases, job postings and UN releases; this excerpt adds per-author productivity estimates and a quality outcome. The concern that polish stops indicating merit matches Does polished writing actually signal better quality work?, though the two draw on different evidence: the review cites evaluator studies, while this excerpt reads the complexity-quality relation directly off manuscript data. The diversity finding runs against Do large language models narrow human expression and thought?, which argues from training statistics without new measurement. Here the measured direction is the opposite: adopters cite more widely. The authors' own blind spot, that the abstract-based detector "almost certainly fails to detect use by authors who heavily edit LLM-assisted text," points toward a different kind of evidence, process data, as in Can process data distinguish AI delegation from ordinary collaboration?, which does not depend on the final text.

What the excerpt does not establish is causation. The authors say so: the detector "relies on abstracts rather than full text," cannot identify which co-author used an LLM, and the adoption timing may be "endogenous to productivity." The excerpt also leaves the quality measure undescribed in the passages shown; it refers to "quality assessments" across complexity bands without explaining them. The 28K peer review reports and 246M online accesses named in the abstract are not developed in the body. The data also predate reasoning models and deep research tools, which the authors flag as a snapshot. The defensible reading is a measured association among detected LLM-assisted preprints. That is strong enough to support the authors' call for journals, funders and tenure committees to rethink writing quality as a proxy for merit, but not enough to say LLMs caused either the productivity gain or the inversion.

Inquiring lines that read this note 5

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How can we detect and account for LLM involvement in academic writing? Do restrictions on reviewer LLM use actually shape peer review behavior? Can AI systems perform peer review as effectively as humans?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 93 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

LLM-adopting scientists produce 23.7-89.3% more papers and LLM-assisted manuscripts invert the link between complexity and quality — across 2.1M preprints