Automation Bias in AI-Assisted Medical Decision-Making under Time Pressure in Computational Pathology

Paper · arXiv 2411.00998 · Published November 1, 2024
Domain Specialization in LLMs

Abstract. Artificial intelligence (AI)-based clinical decision support systems (CDSS) promise to enhance diagnostic accuracy and efficiency in computational pathology. However, human-AI collaboration might introduce automation bias, where users uncritically follow automated cues. This bias may worsen when time pressure strains practitioners’ cognitive resources. We quantified automation bias by measuring the adoption of negative system consultations and examined the role of time pressure in a web-based experiment, where trained pathology experts (n=28) estimated tumor cell percentages. Our results indicate that while AI integration led to a statistically significant increase in overall performance, it also resulted in a 7% automation bias rate, where initially correct evaluations were overturned by erroneous AI advice. Conversely, time pressure did not exacerbate automation bias occurrence, but appeared to increase its severity, evidenced by heightened reliance on the system’s negative consultations and subsequent performance decline. These findings highlight potential risks of AI use in healthcare.

Introduction. Driven by the success of deep learning algorithms in medical imaging, computational pathology seeks to augment practitioner capabilities in areas traditionally challenging for humans like quantitative image analysis tasks e.g., manual biomarker scoring. Given the safety-critical nature of healthcare and the complexity of legal liability for machine misjudgments, clinical decision support systems (CDSS) allow the full benefits of artificial intelligence (AI) integration, including improved diagnostic accuracy and increased efficiency, while keeping the responsibility for the final diagnosis with the medical expert. However, the necessity for practitioner oversight harbors the risk of introducing a new set of challenges, as the mere presence of AI advice could trigger or amplify cognitive biases, which are systematic patterns of deviation from rationality in judgment, such as automation bias (AB). AB refers to the tendency to treat automated cues as infallible, following them unquestioningly instead of vigilant information seeking. This leads to errors when issues go unnoticed because the system fails to detect them (omission errors) or when incorrect automation output is uncritically adopted (commission errors) [1]. Environmental factors like time pressure, ubiquitously present in routine pathology, can place strain on cognitive resources, resulting in heuristic-based usage of decision support systems (DSS) or even automation overdependence [1]. In essence, time stress may amplify both the frequency and magnitude of AB. While most

Related work. While existing research typically evaluates overall acceptance of false AI advice, following Goddard et al. [1] we will measure AB through commission errors, where a previously correct independent evaluation is overturned by incorrect AI guidance (negative consultation), allowing us to isolate AB from other cognitive biases for more precise quantification. Research on the impact of time pressure on AB presents conflicting results, suggesting its effects may be context-dependent: some work demonstrates increased automation dependence under time constraints [2], while others report reduced system reliance in time critical situations [3].

Method. In this paper, for the first time, we quantify AB in human-AI collaboration within computational pathology and examine the additional influence of time pressure through a web-based experiment conducted with non-paid expert users (pathologists, pathology residents, and non-physician pathology staff) independently verified by us and recruited from our professional network, based on prior research collaborations.

Study task. Our study task focused on estimating tumor cell percentage (TCP) on hematoxylin and eosin (H&E)-stained tissue slides, defined as the ratio of neoplastic cells to the total cell count. Typically represented by a single percentage value or specified range, TCP is estimated briefly through visual assessment in clinical practice and is thus suitable to be performed under time constraints. Due to its innate complexity, stemming from the variability of clinical specimen, lack of standardized protocols and absence of formal training, TCP is widely acknowledged as being prone to substantial inter-observer variability [4], rendering it susceptible to cognitive bias manifestation. TCP estimation is commonly performed in routine pathology, as certain molecular tests require a specific level of tumor DNA to ensure robustness and proper interpretation of assay results [4]. Should samples falling below the threshold for a given molecular test be estimated to be above it, unwarranted confidence in the assay result may be fostered, leading to misguided treatment plans and compromised patient care [4]. Materials. For the TCP estimation, we provided participants with a selection of 23 tissue patches, presenting a broad spectrum of tumor cellularity and tissue types at different magnifications. Of these images, three were allocated for a training session, while the remaining 20 were designated for the main experiment. The study material1 was sourced from three openly available datasets, each featuring image patches and dense annotations of various cell types, including tumor cells: the BreCaHad dataset [5], the dataset from Frei et al.’s publication on tumor cell fraction scoring [6], and the BreastPathQ dataset [7]. A standard object detection approach based on the FCOS [8] architecture was utilized to detect tumor- and other cells, from which the TCP for each Based on the literature review the following hypotheses were established: H1: The introduction of AI support, compared to the baseline (no AI), will lead to the emergence of new errors in the form of measurable AB (negative consultations). H2: The presence of time constraints will increase both the frequency and severity of AB, as delineated in H1. The study employed a 2 × 2 factorial, within-subject design with two independent variables (IVs): inclusion of AI (yes/no) and presence of time pressure (yes/no). Automation bias. The occurrence of AB constitutes the key dependent variable (DV). For this, instances of negative consultations, where an initially correct assessment in the baseline treatment (EstB) was altered to a false evaluation in the second round (EstAI) after exposure to an incorrect AI prediction (PredAI), were quantified. To classify the correctness of TCP estimates (ranging from 0-100%), we applied the 25% threshold required for direct DNA sequencing [4]. Assessments were deemed correct if they aligned with the ground truth (GT) in relation to this threshold. For the AB analysis, only dataset entries, where the independent estimate matches the GT, while both the AI recommendation and subsequent AI-aided assessment are erroneous, were considered. After filtering, the remaining data tuples were compared to the total number of AIassisted TCP assessments to derive the AB occurrence rate. Performance. Participant performance, calculated as the mean absolute deviation of TCP estimates from the GT (|DevGT|) per condition, is the second dependent variable.

Discussion. Overall, integration of AI was found to improve participant performance. This ameliorating effect may also explain why the notable performance decline observed under time pressure in the baseline condition was less pronounced during AI-aided evaluations. Additionally, TP appeared to increase alignment with AI advice, likely due to the strain on cognitive resources under stress, which could be considered beneficial when the AI is accurate, but detrimental when the system errs. Interestingly, despite AI’s potential, we recorded that pathologists were largely unwilling to adopt model recommendations, that contradicted their initial judgments, regardless of the correctness of the system’s output. This suggests that AB may not be the primary cognitive bias in AI-aided medical decision-making. Nonetheless, by quantifying the number of accepted negative consultations, we determined an AB occurrence rate of approximately 7% in its "purest" form. This aligns with existing empirical research on AB [1], reporting negative consultation acceptance rates of 6% to 11%. As our findings suggest that interaction with AI can indeed induce AB, leading to errors, that would not have occurred in the absence of system guidance, hypothesis 1 is fully accepted. Contrary to the expectations outlined in H2, the introduction of time pressure did not affect the occurrence frequency of AB. This could be attributed to our modest sample of AB incidents, which might not have been sufficient to capture the effect in question, as well as the nature of our TP simulation. In practice, deadlines manifest in form of volumes of specimen to be assessed rather than individual countdowns, in turn, participants’ reactions may not fully reflect their responses under real clinical time constraints. Consistent with our general observations, time pressure appeared to increase dependence on negative system consultations. This heightened alignment with erroneous AI output might have exacerbated the severity of AB, mirrored in the pronounced performance decline, beyond what would be expected from the negative influence of TP on performance alone, as seen in the general analysis. In summary, while we did not observe significant effects of time strain on AB occurrence, our results indicate that its severity worsens under time stress. Thus, we partially accept hypothesis 2. It has to be acknowledged, that due to the limited availability of expert participants, the study was conducted on a modest sample (n=28). Although a withinsubject design was adopted to augment statistical power, the effects showcased may be under-/over-represented due to sample variability compared to the target population. Moreover, interface design choices such as the omission of clinical background information may have diminished task realism and altered participant behavior, prompting them to approach the study with less diligence as their everyday examinations, potentially impacting observable AB rates. Despite this, the findings serve as a valuable step towards a more holistic understanding of AB in AI-aided medical decision-making and its influencing factors. By highlighting the risk of cognitive biases, we aim to support the safe integration of AI in critical fields like healthcare. Future work could explore the effectiveness of debiasing strategies in minimizing AI-induced cognitive biases such as AB, specifically evaluating their utility under time stress.

Conclusion. Overall, integration of AI was found to improve participant performance. This ameliorating effect may also explain why the notable performance decline observed under time pressure in the baseline condition was less pronounced during AI-aided evaluations. Additionally, TP appeared to increase alignment with AI advice, likely due to the strain on cognitive resources under stress, which could be considered beneficial when the AI is accurate, but detrimental when the system errs. Interestingly, despite AI’s potential, we recorded that pathologists were largely unwilling to adopt model recommendations, that contradicted their initial judgments, regardless of the correctness of the system’s output. This suggests that AB may not be the primary cognitive bias in AI-aided medical decision-making. Nonetheless, by quantifying the number of accepted negative consultations, we determined an AB occurrence rate of approximately 7% in its "purest" form. This aligns with existing empirical research on AB [1], reporting negative consultation acceptance rates of 6% to 11%. As our findings suggest that interaction with AI can indeed induce AB, leading to errors, that would not have occurred in the absence of system guidance, hypothesis 1 is fully accepted. Contrary to the expectations outlined in H2, the introduction of time pressure did not affect the occurrence frequency of AB. This could be attributed to our modest sample of AB incidents, which might not have been sufficient to capture the effect in question, as well as the nature of our TP simulation. In practice, deadlines manifest in form of volumes of specimen to be assessed rather

Limitations. Consistent with our general observations, time pressure appeared to increase dependence on negative system consultations. This heightened alignment with erroneous AI output might have exacerbated the severity of AB, mirrored in the pronounced performance decline, beyond what would be expected from the negative influence of TP on performance alone, as seen in the general analysis. In summary, while we did not observe significant effects of time strain on AB occurrence, our results indicate that its severity worsens under time stress. Thus, we partially accept hypothesis 2. It has to be acknowledged, that due to the limited availability of expert participants, the study was conducted on a modest sample (n=28). Although a withinsubject design was adopted to augment statistical power, the effects showcased may be under-/over-represented due to sample variability compared to the target population. Moreover, interface design choices such as the omission of clinical background information may have diminished task realism and altered participant behavior, prompting them to approach the study with less diligence as their everyday examinations, potentially impacting observable AB rates.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do clinicians calibrate trust in AI medical recommendations? How do users confuse explanation quality with actual system accuracy? How do curriculum design and feedback approaches affect model learning?