Text, images or voice, people spot AI at close to coin-flip odds; the situation may matter more than the format.
Which modality is easiest for humans to detect as AI?
This explores whether people are better at spotting AI in some formats (text, images, voice) than others, and what actually makes AI content detectable to a human at all.
This explores whether some kinds of AI output, such as writing, pictures or voices, are easier for people to catch than others. The short answer from the corpus is that none of them is easy. A review of 30 studies covering text, images and voice found that human accuracy sits close to a coin flip in all three. It has also not improved as AI output has become more realistic Can people reliably spot content made by AI?. The collection doesn't have a head-to-head ranking of the modalities, so it can't name a winner. What it does show is that 'which modality' may be the wrong question.
The situation you're in seems to matter more than the medium. In a 'displaced' Turing test, people who only read transcripts of conversations did worse than chance at telling AI from humans. People who could question the other party in real time kept a small edge Can humans detect AI by passively reading its text?. Being able to ask follow-up questions helps. Simply consuming content, which is how most of us meet AI text, images and audio, takes that advantage away. Our default doesn't help either. In one experiment, people rated unlabeled AI-assisted emails as highly as human-written ones and only became skeptical once the AI involvement was disclosed Do readers trust unlabeled AI-written messages as much as human ones?. People tend to trust content unless they are told otherwise.
The less obvious finding is that AI text is not actually hard to tell apart. It is only hard for people. AI writing differs measurably from human writing in vocabulary variety, yet trained linguists still can't spot it, and newer models drift further from human patterns while becoming harder for people to catch Can humans detect AI text if machines can measure it?. Simple, transparent features catch AI-written Reddit arguments with 99% accuracy. The tells include a habit of mirroring the prompt and a polished, 'textbook' argument style Can simple linguistic features detect AI-written arguments?. In fiction, the giveaway lies deeper than style: choices about how characters act and how time is ordered separate AI stories from human ones with 93% accuracy, even after all stylistic cues are removed Can AI stories be detected without analyzing writing style?. The signal is there. It just sits in patterns that human intuition doesn't pick up.
Voice deserves a separate note. The corpus doesn't measure voice detection directly, but it does show that one strong cue, such as a voice, is enough to make people respond to an AI as a social presence, while a pile of weaker cues is not Do more social cues always make AI feel more present?. That suggests voice may be the format where people are least on guard rather than the easiest to catch, though the collection hasn't tested that. If you want to go further, look at the gap between machine detection and human perception. Detecting AI is turning into a job for instruments rather than instinct, and that changes what it means to tell readers or listeners that content is 'AI-generated.'
Sources 7 notes
A 30-study systematic review found that humans cannot reliably distinguish AI-generated from human-created content across text, image, and voice modalities. Accuracy generally clusters around chance and has not kept pace with improvements in AI realism.
The displaced Turing test shows that both human and AI judges reading transcripts performed below chance accuracy, while interactive interrogators retained marginal detection ability. The adaptive advantage of real-time questioning collapses entirely in passive consumption.
In a preregistered experiment (N=647), recipients rated unlabeled AI-assisted emails indistinguishably from human-written ones. Only explicit AI disclosure triggered strong skepticism. Recipients appear to default to trust rather than suspicion when origin is unrevealed.
LLM-generated text differs significantly on six lexical diversity dimensions, confirmed through statistical analysis across multiple models. Yet human judges, including trained linguists, cannot reliably detect these differences—and newer models diverge further while becoming harder to spot.
General linguistic features combined with argument-quality measures achieved 99% accuracy detecting LLM-generated counter-arguments on r/ChangeMyView, matching heavyweight neural detectors while remaining computationally cheap and transparent. LLMs produce detectable stylistic signatures: accommodation to prompts and textbook-quality argument markers that humans don't replicate.
Show all 7 sources
StoryScope achieved 93.2% accuracy separating AI from human fiction using only discourse-level features like character agency and chronological structure, retaining 97% of performance while eliminating stylistic cues. These structural choices resist humanization because they require rewrites, not surface edits.
Research shows individual primary cues like voice or appearance are sufficient to evoke social-actor presence, while multiple secondary cues cannot. Quality of cues matters more than quantity in driving social responses.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Is it Cake or is it AI? A Systematic Review of Human Uncertainty in Distinguishing Generative Artificial Intelligence Content
- The human-authorship halo: attribution bias in literary style evaluation by humans and AI
- Do LLMs produce texts with "human-like" lexical diversity?
- Linguistic markers of inherently false AI communication and intentionally false human communication: Evidence from hotel reviews
- AI Argues Differently: Distinct Argumentative and Linguistic Patterns of LLMs in Persuasive Contexts
- Measuring AI "Slop" in Text
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- Understanding Reader Perception Shifts upon Disclosure of AI Authorship