Why does a detector trained on one AI model's writing miss text from a different model, and what has it actually learned?
Why do detectors trained on one GPT model fail against another?
This explores why a tool built to spot text written by one AI model (say, GPT-3.5) often stops working when the text comes from a different model (say, GPT-4 or a non-OpenAI model), and what that says about what detectors are actually learning.
This explores why a detector trained to spot one model's writing often misses text from another model, and what that says about what the detector learned. One caveat first: none of the twelve retrieved notes study AI-text detectors, cross-model generalization, or watermarking. What follows is a set of adjacent ideas that suggest a likely answer. It isn't an answer the corpus directly supports, and the gap is worth flagging for the collection.
The most useful idea comes from work that treats an LLM as a probability machine. It shows that a model's behavior is shaped by which outputs its training made likely or unlikely Can we predict where language models will fail?. Two models trained on different data with different methods end up with different maps of what counts as probable. They favor different words, phrasings and sentence rhythms. A detector trained on one model's output may learn that model's particular habits rather than anything common to all machine text. When the generator changes, those habits change too, and the detector has nothing left to go on.
Research on AI judges points the same way from another direction. A single large model used as an evaluator shows intra-model bias, meaning it reacts in a skewed way to text that resembles its own family's style. A panel of smaller judges from different model families turns out fairer and more reliable than any single judge, and no one judge was best everywhere Can a panel of smaller judges outperform one large judge?. Detection likely works the same way. A classifier built around one model family probably picks up that family's quirks, and training on varied sources is the obvious fix. That carries over by analogy, though; the corpus doesn't test it.
It also helps to think of a detector as a discriminator that only makes sense next to the generator it was built against. The Consensus Game frames generation and judgment as two sides that have to converge on a shared view Can generative and discriminative models reach agreement?. A detector has no such shared view with a model it never saw. What it learned was the line between human text and one particular generator, not a general line between human and machine. A related finding may make things harder over time. Models trained on many varied sources tend to settle on smoothed-out, consensus-style output Can models trained on many imperfect experts outperform everyone?, and as models converge in that way, the specific fingerprints a detector depends on may fade.
The main takeaway: a detector probably learns to recognize a particular model, not AI writing as a category. That's why switching models breaks it. For direct evidence on detector transfer, watermarking, or detectors that work across many models, you'll need sources outside this collection for now.
Sources 4 notes
By framing LLMs as autoregressive probability machines, researchers predicted tasks with low-probability target responses would be systematically harder, even when logically simple. Experiments confirmed predictions like backwards alphabet and letter counting.
PoLL (Panel of LLM evaluators) using multiple smaller models from disjoint families outperforms single large judges, reduces intra-model bias, and costs over 7× less. Across three settings and six datasets, no single judge was best everywhere, but diverse panels performed consistently well.
The Consensus Game frames decoding as a signaling game where generator and discriminator must agree on answers. Equilibrium-Ranking finds their joint policy, enabling 7B models to match 540B model performance without fine-tuning.
Models trained on diverse experts converge on consensus behavior that outperforms individuals. Low-temperature sampling concentrates outputs on this majority-voted consensus, denoising uncorrelated biases and errors across the training set.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Transcendence: Generative Models Can Outperform The Experts That Train Them
- The Consensus Game: Language Model Generation via Equilibrium Search
- Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- Large Language Diffusion Models
- Large Language Model Reasoning Failures
- Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
- Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks