INQUIRING LINE

How often can an AI-writing detector wrongly flag a human before it actually hurts honest writers, and where's that line?

What false-positive rate would indicate the classifier harms legitimate human writers?

This explores how often an AI-text detector can wrongly flag human writing before it starts doing real damage to people who wrote their own work, and what the corpus says about where that line sits.


This explores how often an AI-text detector can wrongly flag human-written text before it starts hurting honest writers. The direct answer is that the corpus doesn't give a numeric threshold. No note here says that 1% is fine and 5% is harmful. What it does offer is more useful: reasons why the acceptable rate is probably much lower than intuition suggests, and why the rate alone is the wrong thing to measure.

Start with what a false flag actually costs. When identical passages carry different authorship labels, human judges rate the 'human' version 13.7 points higher, and AI evaluators show a bias two and a half times stronger, at 34.3 points Do authorship labels bias how we judge literary quality?. So a wrong 'AI-written' label doesn't just add a footnote. It can shift how the work is graded or judged by a large margin. Compare that with honest disclosure: when writers volunteer that they used AI, ratings fall by less than 0.15 points on a 7-point scale Does disclosing AI assistance make readers trust articles less?. Being falsely accused seems to cost far more than actually admitting AI help. That gap is a strong argument that even a small false-positive rate imposes costs out of proportion to its size.

The harm also isn't spread evenly. Detectors tend to react to writing style rather than to who actually wrote the text. Fake-news detectors, for example, mistake LLM-like phrasing for deception: they flag truthful AI text as fake and wave through human-written disinformation Why do fake news detectors flag AI-generated truthful content?. An AI-authorship classifier built the same way will flag humans whose natural style happens to resemble model output. Those false positives cluster on particular writers instead of falling at random. Research on human accusers finds the same pattern. Comments that people accused of being AI-written lacked any features that actually separate AI text from human text. The accusations worked as gatekeeping, and the injustice landed on the human writer Do unfounded AI accusations harm human writers instead?.

Meanwhile, the people a detector is meant to catch may slip through. One paper claims that heavily rewritten AI text evades detectors. That claim hasn't actually been tested against detectors yet Do rewrites that hide authorship also fool AI detectors?. Still, it points to an uncomfortable possibility: a detector that misses careful AI users and flags stylistically 'smooth' humans gets worse precisely where it matters. A low overall false-positive rate can hide a high rate for the writers who are flagged most often.

The deeper takeaway is a reframe. Some researchers argue that the useful question isn't 'did a machine write this?' but 'is this text low-quality slop?', a judgment that applies to human and machine writing alike Can we judge text quality without knowing who wrote it?. On that view, the harm threshold is less a percentage to tune than a sign that the classifier is answering the wrong question. If you want to set an actual number, the corpus suggests testing false positives separately for each writer group and style, and weighing them against how much a 'probably AI' label shifts a reader's judgment.


Sources 6 notes

Do authorship labels bias how we judge literary quality?

Human judges rated identical passages 13.7 percentage points higher when labeled human-authored; AI models showed a 2.5-fold stronger bias at 34.3 points. The effect persists across AI architectures, suggesting evaluators respond to provenance cues rather than text quality alone.

Does disclosing AI assistance make readers trust articles less?

Both human raters (n=1,970) and LLM raters (n=2,520) scored an identical news article lower when it included an AI disclosure statement, but the penalty was small—less than 0.15 points on a 7-point scale.

Why do fake news detectors flag AI-generated truthful content?

Fake news detectors flag LLM-generated content as fake while misclassifying human-written disinformation as genuine. The bias arises because detectors trained on human deception patterns mistake AI's distinct linguistic style for falsity, not because they evaluate veracity.

Do unfounded AI accusations harm human writers instead?

Accused comments lack features that distinguish AI text from human writing, suggesting accusations function as gatekeeping rather than detection. This inverts the AI-as-perpetrator framing, placing harm at the receiving side through reader skepticism.

Do rewrites that hide authorship also fool AI detectors?

The paper asserts that rewritten messages evade AI-text detectors but provides no detector experiments, only attribution results showing stylistic convergence. The double erasure claim needs direct empirical testing.

Show all 6 sources
Can we judge text quality without knowing who wrote it?

Research distinguishes slop—a quality assessment based on coherence and relevance—from AI-text detection, which identifies authorship origin. The framework applies equally to human and machine-written texts, separating what a text reads like from who produced it.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.