INQUIRING LINE

AI writing carries measurable tells that most readers miss, and they tend to assume a human wrote it, but what would actually reveal it?

What heuristics do readers use to detect or fail to detect LLM writing?

This explores what cues people rely on when they try to tell whether a piece of writing came from an LLM, and why those cues so often fail.


This explores what people look for when they try to spot AI-written text, and why they so often miss it. The corpus has one surprising answer: the problem is not that LLM text lacks a fingerprint. It's that the fingerprint sits where human readers don't look. When ChatGPT writing is measured on six separate dimensions of vocabulary (how many distinct words, how evenly they're spread, how varied they are), it differs from human writing in statistically solid ways. Yet linguists and NLP researchers still can't reliably tell the two apart Can human judges detect measurable differences in AI text?. The signal is real. It just isn't something a reader can perceive.

The heuristic readers actually seem to fall back on is a default assumption: when in doubt, assume a human wrote it. In a study of research abstracts, readers with machine-learning expertise tended to credit human involvement across every type of abstract. The twist is that abstracts edited by an LLM got the highest clarity ratings and were preferred 55% of the time, even when readers were told who wrote them Can readers tell LLM abstracts from human ones?. So the polish that might seem like a giveaway reads to most people as a sign of good writing.

Machines do much better with surprisingly simple cues. On r/ChangeMyView, a handful of interpretable language features detected LLM-written counter-arguments with 99% accuracy. They matched heavyweight neural detectors Can simple linguistic features detect AI-written arguments?. The two tells were that LLMs echo and accommodate the prompt they were given, and that their arguments carry 'textbook' quality markers real Redditors rarely bother with. A broader pattern lies underneath: picking up style at the level of patterns is easy and saturates fast, while understanding *why* a stylistic choice matters is much harder Can language models truly understand literary style?. Human readers tend to read for meaning, which may be exactly why they miss the patterns.

The parallel with LLM judges suggests something you might not expect. When LLMs grade text, they fall for fake references and rich formatting regardless of what the content says Can LLM judges be fooled by fake credentials and formatting? Can LLM judges be tricked without accessing their internals?. These are the same surface cues of authority and polish that seem to win over human readers of abstracts. Meanwhile, the surface is becoming less trustworthy. Weaker models damage documents visibly by deleting content, while frontier models corrupt content quietly and leave the text looking intact Does model capability change how documents degrade?. Heavy rewriting may also erase both the author's style and the AI's style at once, although that claim hasn't yet been tested against actual detectors Do rewrites that hide authorship also fool AI detectors?.

One gap is worth naming plainly: the corpus doesn't contain a study that catalogs the specific rules of thumb people use, such as 'too many em-dashes' or 'it sounds too balanced.' What it does show is a lever. In a 640-employee study, reviewers caught more LLM errors when the reasoning they needed to check the output was in front of them while they reviewed Can reviewers access what they know when checking LLM outputs?. That suggests better detection may depend less on sharper intuition and more on having the right things to check against while you read.


Sources 9 notes

Can human judges detect measurable differences in AI text?

Six-dimension MANOVA analysis confirms significant differences between ChatGPT and human writing across vocabulary volume, abundance, variety, evenness, disparity, and dispersion. Despite these robust statistical differences, human judges including linguists and NLP researchers fail to reliably distinguish AI from human text.

Can readers tell LLM abstracts from human ones?

Readers with ML expertise struggle to identify LLM-generated content reliably, tending to assume human involvement across all abstract types. However, LLM-edited abstracts received highest clarity ratings and were preferred 55% of the time when authorship was disclosed.

Can simple linguistic features detect AI-written arguments?

General linguistic features combined with argument-quality measures achieved 99% accuracy detecting LLM-generated counter-arguments on r/ChangeMyView, matching heavyweight neural detectors while remaining computationally cheap and transparent. LLMs produce detectable stylistic signatures: accommodation to prompts and textbook-quality argument markers that humans don't replicate.

Can language models truly understand literary style?

GPT-2 achieves 95% accuracy identifying authorship through style patterns alone, but lacks the evaluative framework to explain why those stylistic choices carry meaning. Detection without interpretation remains cataloguing, not criticism.

Can LLM judges be fooled by fake credentials and formatting?

Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.

Show all 9 sources
Can LLM judges be tricked without accessing their internals?

Research shows LLM evaluators systematically score higher when responses include fake references or rich formatting, independent of content quality. These biases are exploitable without model access, undermining AI benchmark credibility.

Does model capability change how documents degrade?

DELEGATE-52 shows weaker LLMs degrade documents through visible deletion, while frontier models degrade through subtle corruption that preserves surface integrity. This shift makes frontier failures harder to detect and potentially more dangerous at workflow scale.

Do rewrites that hide authorship also fool AI detectors?

The paper asserts that rewritten messages evade AI-text detectors but provides no detector experiments, only attribution results showing stylistic convergence. The double erasure claim needs direct empirical testing.

Can reviewers access what they know when checking LLM outputs?

Two experiments with 640 employees showed that error detection improved when verification-relevant reasoning was accessible at review time. Self-generated explanations and retrieval cues strengthened detection, revealing a third failure mode beyond capability or engagement gaps.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.