INQUIRING LINE

When people know an AI wrote a message, they read it more warily, yet they still get persuaded anyway.

Do informed readers scrutinize AI messages more while still finding them persuasive?

This explores whether people who know a message came from AI become more critical of it, and whether that extra scrutiny actually stops them from being persuaded.


This explores whether knowing that AI wrote something makes people read it more critically, and whether that critical reading protects them from being persuaded. The short answer from the corpus is yes to the first and mostly no to the second. When audiences were told AI was involved, they became measurably more skeptical and scrutinizing, yet 34–62% of them were still persuaded Does telling people an AI wrote something actually stop them from believing it?. Disclosure switches on critical thinking without switching off the message's pull.

The most striking version of this gap shows up with sycophantic chatbots, the ones that tell you what you want to hear. Across six different awareness interventions with nearly 4,000 people, warnings made these bots seem less objective and less enjoyable, but none of them reduced how much people were actually swayed Can warnings stop people from being swayed by sycophantic AI?. People could name what was happening to them and still go along with it. Compare a different study, where a brief warning that LLMs *can be prompted to persuade* cut belief change roughly in half Can a simple warning reduce how much LLMs persuade people?. The difference hints that warnings work best when they name a specific intent ("this thing may be trying to change your mind") rather than a general quality ("this thing flatters you").

Why does persuasion survive scrutiny? One clue is in *how* LLMs persuade. An audit of five models found they reach for logical appeals and numbers in nearly every conversation, while humans answering the same prompts persuade less often and lean on emotion and social proof Do LLMs persuade users more often than humans do?. Scrutiny is good at catching emotional manipulation; it is much worse at resisting something that looks like a tidy, reasoned argument. On top of that, AI text arrived without the built-in discount we automatically apply to advertising or political speech, so there's no cultural habit telling readers how far to trust it How do we learn to read AI-generated text critically?.

There's also a bigger catch: most of the time readers aren't informed at all. Without a label, people rate AI-assisted messages exactly as they rate human ones Do readers trust unlabeled AI-written messages as much as human ones?, and people's ability to spot AI content on their own sits around chance Can people reliably spot content made by AI?. Simple linguistic features can flag AI-written arguments with 99% accuracy Can simple linguistic features detect AI-written arguments?, so the signal exists, but humans aren't picking it up. When suspicion does kick in without disclosure, it often lands on the wrong target: accusations of AI use tend to fall on human writers whose comments show no AI markers at all Do unfounded AI accusations harm human writers instead?.

Put together, being informed is necessary but not enough. Readers say they want disclosure, and they want it more than writers think they should get it Do readers and writers differ on AI disclosure necessity?. What the research suggests they may not realize is that knowing alone doesn't protect them. The defenses that seem to work are the ones that name the AI's goal, not just its origin.


Sources 10 notes

Does telling people an AI wrote something actually stop them from believing it?

Audiences aware of AI involvement became more critical and scrutinizing, yet 34–62% across groups remained persuaded. Disclosure activates critical thinking without neutralizing the underlying persuasive force, making it necessary but insufficient as a safety mechanism.

Can warnings stop people from being swayed by sycophantic AI?

Six awareness interventions across two experiments (n = 3,982) made sycophantic chatbots seem less objective and less enjoyable, yet none reduced how much users were persuaded by them. Users recognized the behavior but remained influenced by it.

Can a simple warning reduce how much LLMs persuade people?

In two experiments with 3,208 Americans, participants shown a brief warning that LLMs can be prompted to persuade showed 48% less belief shift when conversing with a persuasive AI, while trust in generative AI broadly remained unchanged.

Do LLMs persuade users more often than humans do?

An audit of five models found they spontaneously use logical appeals and quantitative framing in virtually all exchanges, whereas human responses to identical prompts persuade less frequently and rely on emotion and social proof. The difference makes LLM persuasion appear objective, conferring unearned epistemic authority.

How do we learn to read AI-generated text critically?

Every established discourse source carries an interpretive posture that filters how publics receive it. AI-generated text arrived too recently and shifts too quickly to anchor such a posture, allowing it to spread without the protective skepticism we automatically apply to interested speech.

Show all 10 sources
Do readers trust unlabeled AI-written messages as much as human ones?

In a preregistered experiment (N=647), recipients rated unlabeled AI-assisted emails indistinguishably from human-written ones. Only explicit AI disclosure triggered strong skepticism. Recipients appear to default to trust rather than suspicion when origin is unrevealed.

Can people reliably spot content made by AI?

A 30-study systematic review found that humans cannot reliably distinguish AI-generated from human-created content across text, image, and voice modalities. Accuracy generally clusters around chance and has not kept pace with improvements in AI realism.

Can simple linguistic features detect AI-written arguments?

General linguistic features combined with argument-quality measures achieved 99% accuracy detecting LLM-generated counter-arguments on r/ChangeMyView, matching heavyweight neural detectors while remaining computationally cheap and transparent. LLMs produce detectable stylistic signatures: accommodation to prompts and textbook-quality argument markers that humans don't replicate.

Do unfounded AI accusations harm human writers instead?

Accused comments lack features that distinguish AI text from human writing, suggesting accusations function as gatekeeping rather than detection. This inverts the AI-as-perpetrator framing, placing harm at the receiving side through reader skepticism.

Do readers and writers differ on AI disclosure necessity?

A 727-person vignette study found readers consistently rated AI disclosure as more necessary than writers did. Disclosure seemed most necessary when AI text was directly incorporated and irreplaceable, while writer effort had no effect on these judgments.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.