INQUIRING LINE

Polished AI work impresses at first, until readers spot the polish covering for missing substance and trust the sender less.

Why does polished output make senders seem less capable to recipients?

This explores why people who send polished, AI-generated work sometimes end up judged as less capable by the people who receive it — and the corpus suggests polish alone isn't what does the damage.


This explores why people who send polished, AI-generated work sometimes end up judged as less capable by the people who receive it. The corpus suggests the question's premise needs one correction: polish by itself doesn't lower anyone's opinion of the sender. What lowers it is the moment a recipient notices that polish is standing in for substance. When a BetterUp and Stanford survey asked people about receiving "workslop," about half rated the sender as less creative, capable, and reliable. Forty-two percent trusted them less, and nearly a third were less willing to work with them again Does receiving AI-written work change how we judge the sender?.

What's surprising is that polish usually helps, at least at first. In one preregistered experiment, recipients rated unlabeled AI-assisted emails just as favorably as human-written ones. Their skepticism only kicked in once the AI use was disclosed Do readers trust unlabeled AI-written messages as much as human ones?. Reviews of evaluation studies go further: evaluators often mistook AI-generated documents for human work and rated them higher than real human submissions Does polished writing actually signal better quality work?. So the drop in reputation isn't a reaction to good writing. It's a reaction to finding out the good writing wasn't backed by the sender's own thinking.

Why the backlash is so harsh comes down to a broken shortcut. For a long time, professional-looking work was a fair sign that someone with expertise had made it. Generative AI produces the look without the judgment behind it Does polished AI output trick audiences into trusting it?. The same pattern shows up in model training: models trained to imitate ChatGPT copy its confident, fluent style and fool human evaluators, yet they get no better at being accurate Can imitating ChatGPT fool evaluators into thinking models improved?. When a recipient finds that gap, the polish becomes evidence against the sender. It reads as an attempt to look competent instead of being competent, and it means the recipient now has to do the real work. Frontier models make this worse: they tend to damage documents through subtle errors that leave the surface looking intact, so problems come to light late and after the recipient has already trusted the work Does model capability change how documents degrade?.

The part you may not have expected to want to know is why senders keep doing it. Fluent output fools the person who prompted it too. Users read the ease and quality of AI output as a sign of their own competence, even though they didn't produce it Does processing ease mislead users about their own competence?. Senders therefore aren't knowingly passing off weak work. They honestly believe it's good. Many receivers, for their part, accept fluent output without checking it, a pattern one note calls "cognitive surrender" When do users stop checking whether AI output is actually backed?. That's why workslop circulates so widely before anyone notices.

One more angle: how a message is received depends on more than the text. It also depends on who presents it, how it's framed, and what role the recipient is in. Work on AI explanations makes this point What if XAI is fundamentally a communication problem?. The same polished paragraph can win trust from a stranger who takes it at face value and lose it with a colleague who has to act on it. That suggests the reputational cost of workslop lands hardest inside working relationships, where recipients have both the motive and the context to check.


Sources 9 notes

Does receiving AI-written work change how we judge the sender?

About half of survey respondents who received workslop rated the sender as less creative, capable, and reliable. Forty-two percent viewed them as less trustworthy, and nearly one-third said they'd be less willing to work with them again.

Do readers trust unlabeled AI-written messages as much as human ones?

In a preregistered experiment (N=647), recipients rated unlabeled AI-assisted emails indistinguishably from human-written ones. Only explicit AI disclosure triggered strong skepticism. Recipients appear to default to trust rather than suspicion when origin is unrevealed.

Does polished writing actually signal better quality work?

Studies show evaluators perceived AI-generated documents as both human-written and better quality than human submissions. This suggests rhetorical polish misleads judgment and should not serve as a quality signal in evaluation.

Does polished AI output trick audiences into trusting it?

Generative AI produces visually sophisticated outputs without underlying judgment, leveraging the historical heuristic that professional-looking work signals expert thinking. This substitution is especially risky for less experienced workers who lack domain knowledge to evaluate substance beyond form.

Can imitating ChatGPT fool evaluators into thinking models improved?

Imitation models fool human evaluators by mimicking ChatGPT's confident, fluent style while failing to improve factuality or generalization on novel tasks. The ceiling is set by base model capability, not fine-tuning method—better fundamentals, not shortcuts, drive real improvement.

Show all 9 sources
Does model capability change how documents degrade?

DELEGATE-52 shows weaker LLMs degrade documents through visible deletion, while frontier models degrade through subtle corruption that preserves surface integrity. This shift makes frontier failures harder to detect and potentially more dangerous at workflow scale.

Does processing ease mislead users about their own competence?

High-quality AI output triggers a metacognitive heuristic: users experience fluency as a signal of their own capability, even though they didn't generate it. This self-directed fluency illusion systematically inflates perceived competence because LLMs optimize for fluency regardless of user understanding.

When do users stop checking whether AI output is actually backed?

Users systematically accept AI outputs without verification because checking is costly and fluent output builds false confidence. This receiver-side surrender—measured in studies showing 80% unchallenged adoption—is what enables inflationary token systems to function at scale.

What if XAI is fundamentally a communication problem?

Explanation quality is not intrinsic to the explanation itself but depends on the rhetorical situation: who presents it, how it is framed, and what role the recipient plays. Evaluations that ignore this triad measure only a narrow slice of real-world effectiveness.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.