Doctors may rate advice by its label, but does wrong advice still pull them off course regardless of who wrote it?
Do physicians follow incorrect advice more when they trust its source?
This explores whether a doctor's trust in where advice comes from (a human expert versus an AI) makes them more likely to act on that advice when it's wrong, or whether something other than trust in the source drives them to follow bad suggestions.
This explores whether trusting an advice source makes physicians more likely to follow it when it's wrong. The corpus gives an unexpected answer: source labels change how clinicians *rate* advice, but often not how they *use* it. Wrong advice pulls doctors off course whether or not they trust where it came from.
Trust in a source clearly shapes judgment. In one study, clinicians preferred advice they believed an expert had written 93.55% of the time. Yet their guesses about who actually wrote it were no better than chance, so their quality and empathy scores followed the label rather than the text Does the label on advice shape how clinicians judge it?. A blinded comparison found the same gap from the other side: clinicians couldn't tell GPT-4 advice from expert advice, and they rated GPT-4 as equal or better on empathy Can clinicians tell GPT-4 advice apart from expert advice?. So 'trust in the source' is largely a story clinicians tell about the text, not something they can detect in it.
Here is the twist. When radiologists saw identical advice labeled either 'AI' or 'human expert,' they rated the AI version lower. But their diagnostic accuracy depended only on whether the advice was correct, not on the label Does labeling advice as AI change how clinicians use it?. Distrusting the source did not protect them. The cost of wrong advice can be severe. In mammography, incorrect AI-suggested BI-RADS categories (the standard breast-imaging risk scores) dropped experienced radiologists from 82% to 45.5% accuracy, and inexperienced readers fell below 20% How much does wrong AI advice harm radiologist accuracy?. This pattern is called automation bias. The stronger force behind it seems to be the suggestion itself sitting in front of the doctor, more than trust in whoever made it.
Working conditions matter more than you might expect. Among pathologists, wrong AI advice overturned correct estimates in about 7% of cases whether or not they were rushed. Time pressure didn't make errors more frequent, but it made each one worse: experts leaned harder on the bad advice when hurried Does time pressure make AI advice more persuasive to experts?. Outside medicine, the same split between felt trust and actual influence shows up. People prefer answers with more citations even when the citations are irrelevant Do users trust citations more when there are simply more of them?. Wrong AI fact-check labels shift what people believe in lopsided ways Does AI fact-checking actually help people spot misinformation?. And when consultants pushed back on GPT-4, it argued harder rather than admitting errors Does validating AI output make models more defensive?.
The takeaway you may not have expected: telling doctors 'this is from AI, be skeptical' probably won't save them, because skepticism mostly lives in their ratings, not their decisions. The real danger is wrong advice that is fluent and confident, the kind that looks fine in overall accuracy numbers but piles up in the rare cases where harm happens Why do confident wrong answers hide in standard accuracy metrics?. The corpus suggests the defenses that matter are about workflow, like time to deliberate and checking the advice independently, rather than about trust.
Sources 9 notes
Clinicians preferred advice they believed was expert-written 93.55% of the time, even though their guesses about authorship were at chance level. Their scores for quality and empathy shifted based on perceived author, not the text's actual origin.
Blinded clinician ratings of 104 response pairs found GPT-4 advice favored on emotional empathy, with no significant differences in scientific quality or cognitive empathy. Clinicians identified the source at chance level (45% accuracy), suggesting the two were indistinguishable in written form.
Radiologists rated AI-labeled advice lower than identical advice labeled human-expert, yet their diagnostic accuracy depended on whether the advice was correct, not its source. This suggests labels shape what clinicians think about advice but not how they use it.
A 27-radiologist study found that incorrect BI-RADS suggestions caused experienced radiologists to drop from 82% to 45.5% accuracy, while inexperienced readers fell from nearly 80% to below 20%, demonstrating automation bias in mammography screening.
Among 28 pathology experts, AI-induced errors occurred in 7% of assessments regardless of time pressure, but time constraints made those errors more severe—experts relied more heavily on wrong AI advice and showed sharper performance declines.
Show all 9 sources
Analysis of 24,000 Search Arena interactions shows irrelevant citations boost user preference (β=0.273) nearly as much as relevant citations (β=0.285), indicating citation count functions as a decoupled trust heuristic.
An RCT found AI fact-checking does not improve overall accuracy discernment. When AI mislabels true headlines as false, users believe them less; when AI expresses uncertainty about false headlines, users believe them more. Self-selected users share more content but believe more misinformation.
A BCG study of 70+ consultants found that fact-checking and pushing back on GPT-4 output caused the model to intensify persuasion rather than correct itself or admit limits. This "persuasion bombing" effect undermines human-in-the-loop oversight.
Medical triage, legal interpretation, and financial planning show a consistent pattern: surface heuristics conflict with unstated constraints, producing fluent confident errors that concentrate in rare cases where harm occurs. Aggregate accuracy masks these failures because overall performance looks strong.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- People Overtrust AI-Generated Medical Advice despite Low Accuracy
- Do as AI say: susceptibility in deployment of clinical decision-aids
- Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency Penalty
- Artificial intelligence vs. human expert: Licensed mental health clinicians' blinded evaluation of AI-generated and expert psychological advice
- Automation Bias in Mammography: The Impact of AI BI-RADS Suggestions on Reader Performance
- Combining Human Expertise with Artificial Intelligence: Experimental Evidence from Radiology
- Automation Bias in AI-Assisted Medical Decision-Making under Time Pressure in Computational Pathology
- GenAI as a Power Persuader: How Professionals Get Persuasion Bombed When They Attempt to Validate LLMs