If a person or another AI heavily rewrites machine-written text, do detectors trained on raw AI output still catch it?
Would detectors trained on unaltered AI text catch heavily rewritten versions?
This explores whether AI-text detectors trained on raw, unedited model output still work once someone substantially rewrites that text, whether a person or another model does the rewriting.
This explores whether a detector that learned to spot raw AI output can still catch that text after it has been heavily rewritten. The honest answer from this collection is that nobody here has tested it directly. One paper claims that rewritten messages pull off a 'double erasure': they hide who wrote them and also slip past AI detectors. But its evidence only shows that rewriting makes writing styles converge. It never runs an actual detector on the rewritten text Do rewrites that hide authorship also fool AI detectors?. So the claim is plausible, but it's an assumption, not a finding.
What the corpus does show is that the answer probably depends on which layer of the text a detector looks at. Many detectors rely on surface signals like word choice and vocabulary patterns. Simple, readable features of that kind can be very strong: cheap linguistic measures picked out LLM-written Reddit counter-arguments with 99% accuracy, mostly because models echo the prompt and write in a polished, textbook-argument style Can simple linguistic features detect AI-written arguments?. AI text also differs measurably from human text on six separate measures of vocabulary variety Can human judges detect measurable differences in AI text?. Those surface signals are exactly what a heavy rewrite scrambles, so a detector built on them is the most likely to fail.
The more interesting lead comes from fiction. StoryScope separated AI stories from human ones with 93% accuracy using only narrative choices, such as how much agency characters have and whether events are told in order. It kept almost all of that accuracy after every stylistic cue was removed Can AI stories be detected without analyzing writing style?. These structural choices hold up against 'humanizing' because changing them means rebuilding the piece, not polishing its sentences. That suggests a way to restate the question: rewriting erases a detector's evidence only to the extent that it rewrites the decisions, not just the wording. A paraphrase that leaves the AI's structure in place may still be caught by a detector that looks at structure.
Two more findings show what's at stake. If detectors do fail, people can't serve as a backup. A review of 30 studies found that humans spot AI content at roughly chance level across text, images, and voice Can people reliably spot content made by AI?. Even trained linguists miss differences that statistics pick up easily, and newer models are harder to spot Can humans detect AI text if machines can measure it?. But heavy rewriting may also be rarer than the question assumes. In one study, writers edited AI-drafted paragraphs only 23% of the time, and their edits left the text about 96% similar to the original Do writers actually edit AI-generated text before publishing?. In practice, most AI text reaches readers barely touched. Heavy rewriting is mainly a concern for people deliberately trying to evade detection.
Sources 7 notes
The paper asserts that rewritten messages evade AI-text detectors but provides no detector experiments, only attribution results showing stylistic convergence. The double erasure claim needs direct empirical testing.
General linguistic features combined with argument-quality measures achieved 99% accuracy detecting LLM-generated counter-arguments on r/ChangeMyView, matching heavyweight neural detectors while remaining computationally cheap and transparent. LLMs produce detectable stylistic signatures: accommodation to prompts and textbook-quality argument markers that humans don't replicate.
Six-dimension MANOVA analysis confirms significant differences between ChatGPT and human writing across vocabulary volume, abundance, variety, evenness, disparity, and dispersion. Despite these robust statistical differences, human judges including linguists and NLP researchers fail to reliably distinguish AI from human text.
StoryScope achieved 93.2% accuracy separating AI from human fiction using only discourse-level features like character agency and chronological structure, retaining 97% of performance while eliminating stylistic cues. These structural choices resist humanization because they require rewrites, not surface edits.
A 30-study systematic review found that humans cannot reliably distinguish AI-generated from human-created content across text, image, and voice modalities. Accuracy generally clusters around chance and has not kept pace with improvements in AI realism.
Show all 7 sources
LLM-generated text differs significantly on six lexical diversity dimensions, confirmed through statistical analysis across multiple models. Yet human judges, including trained linguists, cannot reliably detect these differences—and newer models diverge further while becoming harder to spot.
Writers edited AI-generated paragraphs only 23% of the time, with edits averaging 96% similarity to the original. This means AI's opinionated and distorted voice propagates with minimal human filtering before publication.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- The human-authorship halo: attribution bias in literary style evaluation by humans and AI
- AI Argues Differently: Distinct Argumentative and Linguistic Patterns of LLMs in Persuasive Contexts
- Is it Cake or is it AI? A Systematic Review of Human Uncertainty in Distinguishing Generative Artificial Intelligence Content
- Do LLMs produce texts with "human-like" lexical diversity?
- Measuring AI "Slop" in Text
- Understanding Reader Perception Shifts upon Disclosure of AI Authorship
- Linguistic markers of inherently false AI communication and intentionally false human communication: Evidence from hotel reviews