SYNTHESIS NOTE
Topics›Expertise in the Age of AI Content›this note

How much does AI rewriting erase distinctive author voice?

Does heavy AI rewriting weaken the computational signals that identify individual authors? The question matters because it bears on whether AI assistance erases stylistic distinctiveness—a possible cost of polish and consistency.

Synthesis note · 2026-10-06 · sourced from Expertise in the Age of AI Content

Heavy rewriting by an AI writing assistant weakens the signals that let a computer tell one author from another, and the size of the loss depends on the register of the writing. The paper defines the Idiolect Erasure Rate (IER) as "the reduction in authorship-attribution accuracy following AI-assisted rewriting." Under heavy rewriting, the deep attributer LUAR loses 66.5 percentage points on personal blogs and 52.5 on Enron workplace email; the stylometric attributer loses 38.5 points on blogs, 28.7 on Enron and 10.0 on Reuters C50 news. The starting points are strong (LUAR reaches 0.815, 0.713 and 0.710 before rewriting), so the drop is measured from a working attributer. On news, where each journalist writes within a fixed beat, heavy rewriting "produces little deep erasure because topic remains highly predictive of authorship."

The authors read the loss as stylistic convergence, not a change of meaning. Three analyses point that way: rewritten texts become less distinguishable from one another, semantic similarity to the originals stays high, and function-word attribution, which carries little lexical content, also degrades. The loss rises with rewriting intensity. Grammar-only correction has much smaller effects, while even a prompt asking the assistant to preserve the author's voice removes most of the recoverable deep signal. Attribution also fails to accumulate: from original messages it approaches perfect accuracy within several messages, while from rewritten messages it saturates near 50 percent. The choice of attributer matters for a related reason. Shuffling word order drops LUAR from 0.710 to 0.205 but drops MiniLM only from 0.385 to 0.360, so the authors treat LUAR as style-sensitive and MiniLM as a topic-sensitive baseline.

This is the individual-level counterpart to the population result in Does AI writing make all writers sound the same?, which reports AI-assisted paragraphs rated more similar across writers than human-written ones. The paper says it asks a different question: "Rather than measuring homogenization across a population, we quantify the loss of attributable writing style for each author." It is also narrower than Does AI writing assistance change how readers perceive the writer?, which concerns perceived traits, where this paper asks whether the author can still be picked out at all. Its result also bears on Do users truly own the AI-generated content they produce?, since it measures a model's ability to attribute text, not anyone's sense of ownership. The detector claim the paper attaches to this result is taken up in Do rewrites that hide authorship also fool AI detectors?.

The excerpt does not establish what human readers recognize. The authors say IER "measures computational attributability, not human recognition," and list human-recognition studies as future work, so a colleague may still recognize a rewritten email. The setting is closed-set, which they call an upper bound on performance. The corpora are three pre-generative-AI sets of 20 to 50 authors, and the paper evaluates single-pass rewriting by three assistants. The primary rewriter is Qwen2.5-1.5B-Instruct; GPT-4o-mini and Gemini Flash are said to show "comparable levels" of deep erasure, but the excerpt gives no figures for them. Multilingual and longitudinal evaluations are listed as future work. The defensible claim is narrow: in these three registers, heavy rewriting by these assistants removes much of a model's ability to pick out an author's style. Whether that matters to a human reader is outside what was measured.

Inquiring lines that read this note 21

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can readers reliably distinguish AI-written text from human writing? How reliably can humans and AI detectors identify machine-generated text? How do writers navigate authorship and delegation with AI? Can AI systems perform peer review as effectively as humans?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 87 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

Heavy AI rewriting weakens authorship signals far more in blogs and email than in topic-structured news — the Idiolect Erasure Rate measures the drop