Even when AI gives clinicians the right answer, it often doesn't change care. Is the blame distrust, bad weighing, or poor delivery?
Why do clinicians fail to act on correct AI suggestions in real care?
This explores why AI that gives the right answer often doesn't improve what clinicians actually do, and whether the cause is distrust, how the advice gets weighed, or how it's delivered into the workflow.
This explores why correct AI advice so often fails to change clinical outcomes, and where the cause lies: in clinicians' heads, in the advice itself, or in how the advice reaches them. The gap is real. On test cases, an LLM has outperformed hundreds of physicians on differential diagnosis, triage and management Can language models reason better than physicians at diagnosis?. Yet when professional radiologists were given AI predictions, their average performance didn't improve Why don't radiologists benefit from AI predictions?. A capable model doesn't automatically make better care.
The radiology study points to a reasoning error rather than plain distrust. Radiologists gave AI predictions too little weight, and they treated their own judgment as if it were independent of the AI's signal. Weighing two opinions well depends on knowing how much they overlap. Get that wrong and you can't capture the gains of working together, even if the AI is usually right. You might expect the label 'AI' to be the problem, but another study complicates that. Radiologists rated advice lower when it was labeled AI than when the same advice was labeled as coming from a human expert, yet their diagnostic accuracy depended on whether the advice was correct, not on its label Does labeling advice as AI change how clinicians use it?. What clinicians say about AI advice and what they actually do with it can come apart.
Here's the twist that changes the question. The same groups that under-use correct AI can be badly misled by incorrect AI. In mammography, wrong AI-labeled categories dropped experienced radiologists from 82% to 45.5% accuracy, and inexperienced readers fell below 20% How much does wrong AI advice harm radiologist accuracy?. So the core problem isn't that clinicians trust AI too little or too much. They can't reliably tell when the AI is right, and the advice gives them no way to find out. That fits evidence that clinicians couldn't tell GPT-4's written advice from expert advice, guessing its source at chance level Can clinicians tell GPT-4 advice apart from expert advice?. Fluent, plausible text looks the same whether it's right or wrong. A broader account of human-AI interaction describes how confusing a model's fluency with its reliability, plus confirmation bias, can compound into misplaced trust Why do people trust AI outputs they shouldn't?.
The most encouraging evidence suggests the fix is mostly about design. In nearly 40,000 live primary-care visits in Nairobi, clinicians with an LLM safety net made 16% fewer diagnostic errors and 13% fewer treatment errors. The authors credit the asynchronous, workflow-aligned interface and active rollout, not the model's raw ability Can AI safety nets reduce errors in live clinical practice?. The radiology study likewise found that contextual information helped where bare predictions did not Why don't radiologists benefit from AI predictions?. Outside medicine, a historical analysis of agent deployments finds that failures come from missing ecosystem conditions such as trust, social acceptability and fit, not from capability gaps Why do capable AI agents still fail in real deployments?. A therapy study found the same thing from another angle: the same language model reduced distress when delivered through a robot or a structured worksheet, but not through a chatbot Why do robots outperform chatbots in therapy despite identical language models?. The takeaway is that the delivery medium is part of the intervention. A correct suggestion that arrives at the wrong moment, without context, in a format clinicians can't check, may help no more than a wrong one.
Sources 9 notes
In physician-adjudicated vignette experiments and an emergency room study, a large language model outperformed hundreds of physicians on differential diagnosis, reasoning, triage, and clinical management tasks across multiple touchpoints.
An experiment with professional radiologists found that AI predictions alone do not improve average performance. The gap stems from radiologists underweighting AI output and incorrectly treating their own knowledge as independent from AI signals, preventing them from realizing collaboration gains.
Radiologists rated AI-labeled advice lower than identical advice labeled human-expert, yet their diagnostic accuracy depended on whether the advice was correct, not its source. This suggests labels shape what clinicians think about advice but not how they use it.
A 27-radiologist study found that incorrect BI-RADS suggestions caused experienced radiologists to drop from 82% to 45.5% accuracy, while inexperienced readers fell from nearly 80% to below 20%, demonstrating automation bias in mammography screening.
Blinded clinician ratings of 104 response pairs found GPT-4 advice favored on emotional empathy, with no significant differences in scientific quality or cognitive empathy. Clinicians identified the source at chance level (45% accuracy), suggesting the two were indistinguishable in written form.
Show all 9 sources
Rose-Frame identifies map-territory confusion, intuition-reason conflation, and confirmation-bias reinforcement as traps that multiply their distorting effects when they co-occur. Evidence from cross-linguistic overreliance and architectural transformer biases confirms the compounding mechanism operates universally.
In 39,849 clinic visits, clinicians with access to an LLM safety net made 16% fewer diagnostic errors and 13% fewer treatment errors than those without. The authors attribute these gains to asynchronous, interface-optimized design and active deployment strategies, not model capability alone.
Historical analysis from GPS to modern AI shows agent failures consistently result from absent ecosystem conditions—value generation, personalization, trustworthiness, social acceptability, and standardization—rather than capability gaps. Even highly capable systems stall without these five conditions.
A 15-day study with 38 students found that robots and worksheets significantly reduced psychological distress while a chatbot using the same LLM did not. The active ingredient was the medium—social presence and structured format—not language capability.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Do as AI say: susceptibility in deployment of clinical decision-aids
- People Overtrust AI-Generated Medical Advice despite Low Accuracy
- Combining Human Expertise with Artificial Intelligence: Experimental Evidence from Radiology
- Automation Bias in Mammography: The Impact of AI BI-RADS Suggestions on Reader Performance
- Clinical knowledge in LLMs does not translate to human interactions
- Towards Conversational Diagnostic AI
- Superhuman performance of a large language model on the reasoning tasks of a physician
- Towards Accurate Differential Diagnosis with Large Language Models