When a detector claims 94% accuracy, does that mean its flags are usually right, or that it catches most generic content?
Does the 94% figure measure precision or recall of generic content detection?
This explores whether a reported 94% accuracy figure for detecting generic or AI-generated content tells you how often a flagged item really is generic (precision) or how much of the generic content gets caught (recall). The corpus does not contain the source of that 94% figure, so it can't settle the question directly.
This explores whether a reported 94% figure for detecting generic content measures precision (when the detector flags something, how often it's right) or recall (of all the generic content out there, how much the detector catches). The direct answer is that none of the retrieved notes report a 94% detection figure, so the corpus can't say which one it is. The question needs the original paper or report. Check whether it gives a confusion matrix, an F1 score, or plain 'accuracy'. A single headline number is often plain accuracy, and on an imbalanced dataset accuracy can look high while one of the two measures is poor.
The corpus does explain why one detection number shouldn't be trusted on its own. A 30-study review found that people can't reliably tell AI-generated content from human-made content in text, images, or voice. Their accuracy hovers around chance and hasn't improved as AI output has gotten more realistic Can people reliably spot content made by AI?. So a high score like 94% is either a strong automated result or a measure of something narrower than it sounds. Which of the two applies depends on how it was measured.
What counts as 'detected' matters too. In an 81-person study, readers with no cues about where claims came from couldn't separate fluent fabrications from true statements. Their discernment came back only when an interface showed how many claims had been verified Can readers tell truth from fabrication without evidence signals?. Detection is a property of the setup as much as of the detector: what signals are available, and what fraction of the test set is actually generic. Change the mix and precision and recall move in different directions.
Automated judges have their own problem. LLM evaluators can be fooled by fake references and polished formatting Can LLM judges be fooled by fake credentials and formatting?. A detector that scores well on recall against plain generic text could lose precision, or simply miss items, once that text is dressed up with citations and headers. That's a good reason to ask what kind of generic content a 94% figure was tested on.
If you're trying to read the figure, ask three things. Is it accuracy, precision, or recall? What share of the test items were generic? Was the generic content plain or disguised? The corpus can help with how to read the number, but the answer about the number itself has to come from its source.
Sources 3 notes
A 30-study systematic review found that humans cannot reliably distinguish AI-generated from human-created content across text, image, and voice modalities. Accuracy generally clusters around chance and has not kept pace with improvements in AI realism.
In an 81-person study, participants given no provenance cues showed no significant truth discernment (p = .43), falling for fluent hallucinations as readily as ground truth. An idealized Provenance Density interface showing verified claims restored a +4.15 point gap (p < .001).
Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Is it Cake or is it AI? A Systematic Review of Human Uncertainty in Distinguishing Generative Artificial Intelligence Content
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency Penalty
- Humans or LLMs as the Judge? A Study on Judgement Biases
- The human-authorship halo: attribution bias in literary style evaluation by humans and AI
- Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
- References Improve LLM Alignment in Non-Verifiable Domains
- LLM-REVal: Can We Trust LLM Reviewers Yet?