Do as AI say: susceptibility in deployment of clinical decision-aids
Source: Gaube et al., npj Digital Medicine · 2021-02
Artificial intelligence (AI) models for decision support have been developed for clinical settings such as radiology, but little work evaluates the potential impact of such systems. In this study, physicians received chest X-rays and diagnostic advice, some of which was inaccurate, and were asked to evaluate advice quality and make diagnoses. All advice was generated by human experts, but some was labeled as coming from an AI system. As a group, radiologists rated advice as lower quality when it appeared to come from an AI system; physicians with less task-expertise did not. Diagnostic accuracy was significantly worse when participants received inaccurate advice, regardless of the purported source. This work raises important considerations for how advice, AI and non-AI, should be deployed in clinical environments.
The data-intensive nature of healthcare makes it one of the most promising fields for the application of artificial intelligence (AI) and machine learning algorithms1–3. Applications of AI in classifying medical images have demonstrated excellent performance in several tasks, often on par with, or even above, that of human experts4,5. However, it is not clear how to effectively integrate AI tools with human decision-makers; indeed, the few cases where systems have been implemented and studied showed no improved clinical outcomes6,7.
AI systems will only be able to provide real clinical benefit if the physicians using them are able to balance trust and skepticism. If physicians do not trust the technology, they will not use it, but blind trust in the technology can lead to medical errors8–11. The interaction between AI-based clinical decision-support systems and their users is poorly understood, and studies in other domains have garnered inconsistent and complex findings. Reported behaviors include both a skepticism or distrust of algorithmic advice (algorithmic aversion)12–14 and more willingness to adhere to algorithmic advice over human advice (algorithmic appreciation)15,16. Responses can vary depending on the task at hand or the person’s task expertise—for instance, one study found that algorithmic appreciation waned when the participants had high domain expertise15,17. It is therefore important to study how physicians of different expertise levels will perceive and integrate AI-generated advice before such systems are deployed18–20.
The participants were physicians with different levels of task expertise: radiologists (n = 138) were the high-expertise group and physicians trained in internal/emergency medicine (IM/EM, n = 127) were the lower expertise group (because they often review chest X-rays, but have less experience and training than radiologists).
We selected eight cases, each with a chest X-ray, from the open source MIMIC Chest X-ray database20. Participants were provided with the chest X-rays, a short clinical vignette, and diagnostic advice that could be used for their final decisions. They were asked to (1) evaluate the quality of the advice through a series of questions, and (2) make a final diagnosis (see Fig. 1).
Each participant reviewed eight cases. For each case, the physician would see the chest X-ray as well as diagnostic advice, which would either be accurate or inaccurate. The advice was labeled as coming either from an AI system or an experienced radiologist. Participants were then asked to rate the quality of the advice and make a final diagnosis.
We tested whether the advice quality ratings were affected by the independent variables (see Table 1 for statistics). As expected, participants across the medical specialties correctly rated the quality of the advice on average to be lower if the advice given to them was inaccurate (see Fig. 2a). The effect was much stronger among task experts (i.e., radiologists) than non-experts (i.e., IM/EM physicians). We note that only participants with higher task expertise showed algorithmic aversion by rating the quality of advice to be significantly lower when it came from the AI in comparison to the human. The main effects remained constant when controlled for the inter-individual variables among both physician groups (see Supplementary Table 1). The advice quality rating correlated significantly with the confidence in their diagnosis among both task experts r(1102) = 0.43, p < 0.001 and non-experts r(1014), p < 0.001.
We then tested if the participants’ final diagnostic accuracy was affected by the source and/or the accuracy of advice (see Table 2 for statistics). As expected, both participant groups showed a higher diagnostic accuracy when they received accurate advice in comparison to inaccurate advice (see Fig. 3a). Task experts performed 40.10% better and non-experts performed 37.53% better when receiving accurate rather than inaccurate advice. Importantly, the purported source of the advice did not affect participants’ performance (see Fig. 3b). The two main effects and interaction did not change after controlling for the same covariates as above (see Supplementary Table 2). Among the task experts, the two covariates of professional identification (p = 0.042) and years of experience (p = 0.003) were associated with higher diagnostic accuracy, while none of the covariates affected the diagnostic accuracy among non-experts. Both task experts and non-experts had significantly more confidence in their diagnosis when it was accurate (radiology: t = 6.65, p < 0.001; IM/EM: t = 8.43, p < 0.001).
As shown in Fig. 4, we investigated the performance of individual radiologists and IM/EM physicians. Radiologists are better performers (13.04% had perfect accuracy, 2.90% had ≤ 50% accuracy), than IM/EM physicians (3.94% had perfect accuracy, 27.56% had ≤ 50% accuracy). We define clinical susceptibility as the propensity to follow incorrect advice, and we find that 41.73% of IM/EM physicians are susceptible, i.e., they always give the wrong diagnosis with inaccurate advice. This is true only for 27.54% of radiologists. Even among physicians with relatively high overall accuracy, a significant portion are susceptible. On the other hand, some physicians are more critical of incorrect advice: 28.26% of radiologists and 17.32% of IM/EM physicians refuted all incorrect advice they were given. Further analysis by advice source is in Supplementary Fig. 1.
We also looked at participants’ performance on each individual case. As shown in Fig. 5, all cases were impacted by incorrect advice to varying degrees. Case 4 has relatively high average performance under both advice types, for both radiologists and IM/EM physicians. In contrast, Case 6 is more difficult, with generally lower performance under both advice types. Respondents may have misinterpreted the superimposition of ribs as a “pseudo-nodule”25; use of the window level/width and magnifying tools in the DICOM viewer should have given the correct diagnoses of a hiatus hernia, apical pneumothorax, and broken rib. Cases that exploit known weaknesses of X-ray evaluators (Cases 3 and 8) had large gaps between diagnostic accuracy under different advice types (inaccurate vs. accurate). For example, in order to correctly diagnose Cases 3 and 8, respondents would need to be aware that pathology is often missed in the retrocardiac window and lung apices26,27. There may be a particular risk of over-reliance on inaccurate advice for such cases where physicians fail to recognize known pitfalls in X-ray interpretation or to perform additional analyses to address them.
Hospitals are increasingly interested in implementing AI-enabled clinical support systems to improve clinical outcomes, where an automated system may be viewed as a regulated advice giver28. However, how AI-generated advice affects physicians’ diagnostic decision-making in comparison to human-generated advice has been understudied. Our experiments work to build some of this understanding and raise important considerations in the deployment of clinical advice systems.
First, providing diagnostic advice influenced clinical decision-making, whether the advice purportedly came from an AI system or a fellow human. Physicians across expertise levels often failed to dismiss inaccurate advice regardless of its source. In contrast to prior work, we did not find that participants were averse to following algorithmic advice when making their final decision12–14,29. We also did not find evidence of algorithmic appreciation, which is in line with previous research exploring behaviors of people with high domain expertise15,17. Rather, we found a general tendency for participants to agree with advice; this was particularly true for physicians with less task expertise. The provided diagnosis could have engaged cognitive biases, by anchoring participants to a particular diagnosis, and triggering confirmatory hypothesis testing where participants direct their attention towards features consistent with the advice30.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How do clinicians calibrate trust in AI medical recommendations?- Do radiologists' beliefs about AI-assisted performance match their actual outcomes?
- Does showing AI confidence scores reduce radiologist over-reliance on wrong suggestions?
- What safeguards help radiologists maintain independent judgment when using AI assistance?
- Why do clinicians fail to act on correct AI suggestions in real care?
- Why do experts resist AI recommendations that contradict their own judgments?
- How much diagnostic accuracy is gained when physicians receive expert advice?
- Does this colonoscopy finding apply to other medical specialties using AI?
- Why did lay users with AI models fail to match unaided physicians on diagnosis?
- Why do nurses misclassify emergencies differently with misleading AI assistance?
- Does AI change clinician cognition or just increase reliance on predictions?
- Why do radiologists fail to benefit from AI decision support?
- How does trusting wrong AI advice change what medical action people decide to take?
- Can an AI system trained on text consultations handle diagnostic uncertainty in real patient encounters?
- What evidence would prove medical AI actually works in clinics?
- Do patients actually perceive AI as worse at addressing their unique medical needs?
- Why did primary care physicians review only 73% of AI-generated transcripts?
- Can clinicians reliably distinguish high-quality AI advice from low-quality advice by appearance alone?
- Do physicians follow incorrect advice more when they trust its source?
- Can annotation and explanation labels reduce automation bias in clinical settings?
- Do expert physicians also prefer AI-written medical text when it is unlabeled?
- Why does labeling advice as AI from a doctor change how people trust it?
- Can people tell which medical advice is accurate based only on how it reads?