When doctors use AI, does it actually change how they think through a case — or just how much they trust the machine's answer?
Does AI change clinician cognition or just increase reliance on predictions?
This explores whether AI tools actually change how clinicians think through a case (what they consider, how they weigh evidence) or mostly change how much they lean on whatever answer the machine gives.
This explores whether AI changes how clinicians reason through a case, or mostly changes how much they defer to the machine's answer. The short version is that the corpus shows both, and the line between them is blurrier than the question suggests. The strongest evidence of a real change in thinking is in differential diagnosis. When clinicians working through hard NEJM cases had access to an LLM, their top-10 differential accuracy rose from 36% to nearly 52%. The authors credit the gain to the model widening the set of possibilities clinicians considered, not to it handing them the right answer Does LLM assistance help clinicians build better differentials?. That changes what goes into the clinician's reasoning, not just what comes out of it.
When the AI gives a prediction rather than a list of options, the picture looks more like reliance, and badly calibrated reliance. In an ICU simulation, wrong AI predictions hurt nurse performance by 96–120%, while correct ones helped by only 53–67% Do wrong AI predictions hurt more than right ones help?. Averaging gains and losses into one score hides this imbalance. Radiologists show the opposite problem. They tend to underweight AI predictions and wrongly treat their own judgment as independent of the AI's signal, so on average they don't benefit at all Why don't radiologists benefit from AI predictions?. So the issue isn't simply too much reliance. Clinicians struggle to fold a machine's opinion into their own judgment in the right proportion: they take on its errors and fail to collect its value.
A telling detail: what clinicians say about AI and what they do with it come apart. Radiologists rated advice lower when it was labeled as coming from AI, yet their accuracy tracked whether the advice was correct, not where it came from Does labeling advice as AI change how clinicians use it?. Clinicians also can't reliably tell GPT-4's written advice from expert advice; they guessed the source at chance level Can clinicians tell GPT-4 advice apart from expert advice?. Their stated skepticism doesn't seem to act as a filter. The advice gets in either way.
The outcome depends heavily on how the collaboration is set up. HealthBench found that frontier models beat physicians working alone, but physicians using the same model matched or exceeded it Do AI models outperform physicians on health tasks?. The same model produces different results depending on who uses it and how. On the AI side, wrapping a model in a structured diagnostic process raised accuracy and cut costs by 70%, with no change to the model itself Can orchestration strategies boost diagnostic AI without better models?. Read together, these suggest that the workflow, more than the model, decides whether AI expands a clinician's thinking or replaces it. That matters more as models outperform physicians on standalone tasks Can language models reason better than physicians at diagnosis?.
The gap you may not have expected: none of these clinical studies measures what happens to a clinician's own reasoning over months of AI use. They measure outcomes in single sessions. Evidence from outside medicine is worrying. A four-month EEG study found that brain connectivity and memory of one's own work declined as people relied more on an LLM Does AI assistance weaken our brain's ability to think independently?. Research on models themselves offers a parallel: training can raise answer accuracy while the quality of the reasoning behind it gets worse, and accuracy metrics miss the decline Does supervised fine-tuning improve reasoning or just answers?. If the same holds for clinicians, then rising diagnostic accuracy with AI could coexist with weakening independent reasoning, and the current studies aren't designed to catch it.
Sources 10 notes
In a study of 20 clinicians on 302 NEJM cases, those with LLM access achieved 51.7% top-10 accuracy versus 36.1% without it. The authors attribute the gain to the LLM's wider differential scope, making lists more comprehensive.
In an ICU simulation, misleading AI predictions degraded nurse performance by 96–120%, while correct predictions improved it by only 53–67%. This asymmetry was hidden by standard metrics that average gains and losses together.
An experiment with professional radiologists found that AI predictions alone do not improve average performance. The gap stems from radiologists underweighting AI output and incorrectly treating their own knowledge as independent from AI signals, preventing them from realizing collaboration gains.
Radiologists rated AI-labeled advice lower than identical advice labeled human-expert, yet their diagnostic accuracy depended on whether the advice was correct, not its source. This suggests labels shape what clinicians think about advice but not how they use it.
Blinded clinician ratings of 104 response pairs found GPT-4 advice favored on emotional empathy, with no significant differences in scientific quality or cognitive empathy. Clinicians identified the source at chance level (45% accuracy), suggesting the two were indistinguishable in written form.
Show all 10 sources
HealthBench's evaluation of 5,000 multi-turn health conversations found frontier models scored higher than physicians working alone, but physicians matched or exceeded model performance when assisted by that same model, suggesting AI benefits depend on who deploys it.
On 304 NEJM cases, MAI-DxO orchestration achieved 79.9% accuracy at $2,397 per case versus 78.6% at $7,850 for o3 alone. The gains transferred across model families, suggesting the benefit comes from the scaffold, not model weights.
In physician-adjudicated vignette experiments and an emergency room study, a large language model outperformed hundreds of physicians on differential diagnosis, reasoning, triage, and clinical management tasks across multiple touchpoints.
A four-month EEG study of 54 participants found that brain connectivity systematically scaled down with AI reliance—LLM users showed weakest neural engagement, poorest memory retention, and impaired ability to recall their own recent work.
Supervised fine-tuning improves final-answer accuracy on benchmarks but cuts Information Gain by 38.9 percent, meaning models generate correct answers through post-hoc rationalization rather than genuine inferential steps. Standard metrics miss this degradation because they only measure final correctness.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- People Overtrust AI-Generated Medical Advice despite Low Accuracy
- Sequential Diagnosis with Language Models
- Clinical knowledge in LLMs does not translate to human interactions
- Do as AI say: susceptibility in deployment of clinical decision-aids
- Combining Human Expertise with Artificial Intelligence: Experimental Evidence from Radiology
- Towards Conversational Diagnostic AI
- AI-based Clinical Decision Support for Primary Care: A Real-World Study
- Towards Accurate Differential Diagnosis with Large Language Models