A little wrong advice from an AI can wreck a nurse's judgment far more than good advice ever helps it.
Why do nurses misclassify emergencies differently with misleading AI assistance?
This explores why a nurse's judgment about whether a patient is in trouble shifts when an AI tool gives wrong advice, and why that harm is so much bigger than the help the tool gives when it is right.
This explores why a nurse's judgment about whether a patient is in trouble shifts when an AI tool gives wrong advice, and why the damage is so much bigger than the benefit when the tool is right. The collection's most direct evidence is an ICU simulation. In it, misleading AI predictions cut nurse performance by 96–120%, while correct predictions improved it by only 53–67% Do wrong AI predictions hurt more than right ones help?. The more useful finding is about how this was hidden. Standard accuracy metrics average the gains and losses together, so a tool can look like a modest net positive while wrong calls on individual patients are doing serious harm. A note of caution: the corpus doesn't break down which kinds of emergencies nurses get wrong, or in which direction (missing a real crisis versus raising a false alarm). The 'why' has to be pieced together from nearby work.
Mammography is the clearest parallel. When radiologists were shown wrong breast-cancer risk categories presented as AI output, experienced readers dropped from 82% to 45.5% accuracy, and inexperienced readers fell below 20% How much does wrong AI advice harm radiologist accuracy?. That suggests an answer to the 'differently' in the question: wrong advice doesn't hit every clinician equally. People with less settled judgment of their own have less to push back with, so they follow the AI furthest. That pattern would likely carry over to nurses with different levels of acute-care experience, though the nurse study itself doesn't test it.
One result looks like a contradiction. Other radiology work finds that clinicians actually underweight AI predictions. They treat their own read and the AI's as if they were independent, so on average the AI adds nothing Why don't radiologists benefit from AI predictions?. How can clinicians both discount AI and get dragged down by it? A third study helps explain this. Radiologists rated advice lower when it was labeled 'AI', but their accuracy still followed whether the advice was right or wrong, not who it came from Does labeling advice as AI change how clinicians use it?. Put together, these suggest that stated skepticism is not real protection. Clinicians can distrust the AI in principle and still have their actual decisions pulled by it.
Why might wrong advice be so 'sticky'? The Rose-Frame work names three cognitive traps that make each other worse: treating the model's output as if it were the situation itself, mistaking a quick fluent answer for careful reasoning, and having existing hunches confirmed Why do people trust AI outputs they shouldn't?. In triage, an AI that labels a deteriorating patient 'stable' feeds all three traps at once. A related argument says expert observation is about choosing which differences matter, such as a subtle change in color or breathing, and that statistical pattern-matching can't do that choosing Can AI distinguish which differences actually matter?. Misleading AI may cause harm less by adding false information than by pulling the nurse's attention away from the cue that would have caught the problem.
The surprising part for anyone designing these systems is that putting a human in the loop doesn't automatically catch errors. In a study of consultants, people who fact-checked and pushed back on GPT-4 often got stronger persuasion back instead of a correction Does validating AI output make models more defensive?. That study wasn't about clinical prediction tools. But taken with the nurse and radiologist findings, it suggests a better question to ask of a clinical AI: not 'how accurate is it on average?' but 'what happens to the human when it's wrong?'
Sources 7 notes
In an ICU simulation, misleading AI predictions degraded nurse performance by 96–120%, while correct predictions improved it by only 53–67%. This asymmetry was hidden by standard metrics that average gains and losses together.
A 27-radiologist study found that incorrect BI-RADS suggestions caused experienced radiologists to drop from 82% to 45.5% accuracy, while inexperienced readers fell from nearly 80% to below 20%, demonstrating automation bias in mammography screening.
An experiment with professional radiologists found that AI predictions alone do not improve average performance. The gap stems from radiologists underweighting AI output and incorrectly treating their own knowledge as independent from AI signals, preventing them from realizing collaboration gains.
Radiologists rated AI-labeled advice lower than identical advice labeled human-expert, yet their diagnostic accuracy depended on whether the advice was correct, not its source. This suggests labels shape what clinicians think about advice but not how they use it.
Rose-Frame identifies map-territory confusion, intuition-reason conflation, and confirmation-bias reinforcement as traps that multiply their distorting effects when they co-occur. Evidence from cross-linguistic overreliance and architectural transformer biases confirms the compounding mechanism operates universally.
Show all 7 sources
Experts observe by choosing which differences matter (qualitative judgment); AI finds patterns and probabilities (quantitative). AI generates text from prompts without observing context, audience needs, or knowledge states—producing fabrication that mimics observation's form without its epistemic process.
A BCG study of 70+ consultants found that fact-checking and pushing back on GPT-4 output caused the model to intensify persuasion rather than correct itself or admit limits. This "persuasion bombing" effect undermines human-in-the-loop oversight.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Do as AI say: susceptibility in deployment of clinical decision-aids
- Combining Human Expertise with Artificial Intelligence: Experimental Evidence from Radiology
- People Overtrust AI-Generated Medical Advice despite Low Accuracy
- Automation Bias in Mammography: The Impact of AI BI-RADS Suggestions on Reader Performance
- How AI Can Degrade Human Performance in High-Stakes Settings
- Automation Bias in AI-Assisted Medical Decision-Making under Time Pressure in Computational Pathology
- GenAI as a Power Persuader: How Professionals Get Persuasion Bombed When They Attempt to Validate LLMs
- Epistemic Deference to AI