SYNTHESIS NOTE
Topics›Knowledge After the Web›this note

Do wrong AI predictions hurt more than right ones help?

When AI tools give incorrect medical predictions, do they damage clinician performance more severely than correct predictions improve it? This matters for understanding whether averaging test results can hide dangerous asymmetries in AI safety.

Synthesis note · 2026-10-09 · sourced from Knowledge After the Web

Morey, Rayo and Woods (Ohio State), writing in AI Frontiers, report a case study using a method they call Joint Activity Testing, in which "450 nursing students and a dozen licensed nurses each reviewed 10 historical ICU cases, using an AI early warning tool in four configurations." Nurses rated each case 1-10 for how concerned they were about the patient. The result: "When AI predictions were most correct, nurses performed 53% to 67% better than when they worked without AI assistance. However, when AI predictions were most misleading, nurses performed 96% to 120% worse than when they worked without AI assistance." Adding annotations to the AI's predictions "did not significantly alter results."

The authors argue this is not a matter of effort or conscious reliance: "these results do not stem from the nurses' lack of effort or from their consciously offloading decision-making to the AI. Instead, AI assistance appeared to change how nurses think when assessing patients." Nurses and the algorithm looked complementary when tested apart — "the algorithm struggled with cases that nurses without AI handled with ease" — yet when misleading predictions were paired with those same routine cases, nurses "consistently misclassified emergencies as nonemergencies (and vice versa)." Their larger argument is methodological: standard evaluations that "treat AI and humans separately" or "boil down the complex effects of collaboration into a single metric, like average performance" hide this asymmetry, because "average gains can mask rare but severe failures." Neither years of experience nor familiarity with AI tools reliably predicted who would benefit or recover from AI's mistakes.

This matches the shape of two automation-bias findings already in the library while adding a measurement neither has: a matched gain-versus-loss figure within one study and task. How much does wrong AI advice harm radiologist accuracy? shows only the loss side — a wrong AI label cutting accuracy — without a matched measure of how much a correct label helped. Does time pressure make AI advice more persuasive to experts? measures a narrower commission-error rate and names time pressure as the severity driver; Morey, Rayo and Woods instead attribute the swing to AI assistance changing cognition generally, with no time-pressure manipulation in view. All three studies reject effort or face-value judgment as the explanation, but this one is the most explicit that averaging itself, not just human bias, is what conceals the effect, and it proposes testing "a range of challenging cases with varied AI performance" instead of a single aggregate score.

The excerpt gives the headline percentages but not the underlying npj Digital Medicine paper's statistics — how the 450 students and 12 licensed nurses were split in the results, significance tests, or how the "four configurations" differed beyond annotations. The task was a simulation on 10 historical cases rated on a 1-10 scale, not live bedside decisions, and the excerpt does not describe how the "without AI assistance" comparison condition was run. The claim that AI "changed how nurses think" rather than caused offloading rests on the authors' own follow-on analyses, not shown here. If the pattern generalizes past this ICU simulation, it implies that a single aggregate accuracy score cannot certify a human-AI tool as safe for high-stakes deployment, whatever that average score is.

Inquiring lines that read this note 7

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do users confuse explanation quality with actual system accuracy? How do clinicians calibrate trust in AI medical recommendations?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 85 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

Morey, Rayo and Woods find misleading AI predictions hurt nurse performance by 96 to 120 percent — more than correct predictions helped by 53 to 67 percent