Clinicians rated AI medical advice about as sound as expert advice, yet still favored whichever they were told an expert wrote. Do patients?
Do patients show the same bias toward expert-labeled medical advice?
This explores whether patients, like clinicians, rate medical advice more highly when they think an expert wrote it, even when the text is identical or AI-generated.
This explores whether patients, like clinicians, judge medical advice by who they think wrote it rather than by what it says. The short answer is that the collection has no study of patients. What it has is strong evidence about clinicians, plus a parallel finding from creative writing. Together they suggest the bias isn't specific to medicine or to experts, though that is an inference, not a measured result.
The clinician evidence is striking because two findings sit side by side. When clinicians compared GPT-4 advice with expert-written advice without seeing labels, they rated the two about equal on scientific quality. They rated GPT-4 slightly higher on emotional empathy, and guessed the source correctly only 45% of the time, which is chance level Can clinicians tell GPT-4 advice apart from expert advice?. Yet the same clinicians preferred whichever advice they believed came from an expert 93.55% of the time Does the label on advice shape how clinicians judge it?. Their quality and empathy scores moved with their guess about the author, not with the real author. So the 'expert' label was doing work that the text itself wasn't. If trained professionals who can't tell the difference still reward the label, there is little reason to expect patients, who have less ability to check medical content, to resist it better.
The closest thing to a test with ordinary readers comes from literary judgment. Human judges rated identical passages 13.7 percentage points higher when told a human wrote them Do authorship labels bias how we judge literary quality?. That suggests label bias is a general habit of human evaluation, not something that comes with clinical training. The surprise is that AI evaluators showed the same bias about 2.5 times more strongly. This matters for patients, who increasingly get advice that an AI has filtered, summarized or ranked. The bias may not only sit in the patient's head. It may already be built into the systems that decide which advice reaches them.
There is also a reason the label carries so much weight. One note argues that an expert claim gets its force from the expert's reputation, track record and standing, which text alone doesn't carry Can language models distinguish expert arguments from common assumptions?. Seen this way, trusting an 'expert-written' label isn't irrational. It stands in for the social trust people normally rely on. The problem comes when the label is wrong or the stand-in no longer matches who actually wrote the advice.
If you want to think about fixes, one design idea changes the role of AI: instead of handing people an answer they will either defer to or dismiss, the system points out which parts of the input matter and leaves the decision with the human Can AI guidance reduce anchoring bias better than AI decisions?. That work targets anchoring bias rather than authorship labels, but the logic carries over. Advice that shows its reasoning gives readers something to judge besides the byline. An actual patient-side study is still a gap in the collection.
Sources 5 notes
Blinded clinician ratings of 104 response pairs found GPT-4 advice favored on emotional empathy, with no significant differences in scientific quality or cognitive empathy. Clinicians identified the source at chance level (45% accuracy), suggesting the two were indistinguishable in written form.
Clinicians preferred advice they believed was expert-written 93.55% of the time, even though their guesses about authorship were at chance level. Their scores for quality and empathy shifted based on perceived author, not the text's actual origin.
Human judges rated identical passages 13.7 percentage points higher when labeled human-authored; AI models showed a 2.5-fold stronger bias at 34.3 points. The effect persists across AI architectures, suggesting evaluators respond to provenance cues rather than text quality alone.
LLMs lose the social context that gives expert claims their force—reputation, track record, and standing—because they process only text, not the social world where expertise is built and evaluated.
Learning to Guide eliminates anchoring bias and unassisted hard cases by having machines supply interpretive guidance rather than autonomous decisions, keeping responsibility with humans while improving their judgment through enhanced perception.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- People Overtrust AI-Generated Medical Advice despite Low Accuracy
- People Defer to AI Moral Advice, But Not Blindly
- Artificial intelligence vs. human expert: Licensed mental health clinicians' blinded evaluation of AI-generated and expert psychological advice
- Do as AI say: susceptibility in deployment of clinical decision-aids
- The human-authorship halo: attribution bias in literary style evaluation by humans and AI
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- Learning To Guide Human Experts Via Personalized Large Language Models
- Understanding Reader Perception Shifts upon Disclosure of AI Authorship