INQUIRING LINE

Strong scores on rare, puzzle-like medical cases are impressive, but do they predict how an AI helps in a real clinic?

Can LLM performance on zebra cases predict results in routine clinical practice?

This explores whether an LLM's strong showing on rare, puzzle-like diagnostic cases (the 'zebras' found in published case challenges such as those in the New England Journal of Medicine) tells us how much it will help with the everyday work of a real clinic.


This explores whether an LLM's strong showing on rare, puzzle-like diagnostic cases tells us how much it will help in everyday clinical work. The short answer from the corpus is: not much on its own. None of these studies runs the direct test, which would mean taking one model, measuring it on zebra cases, and then checking whether that score predicts its value in routine practice. What the corpus does show is a consistent gap between how a model performs on a case and how much it helps the people and workflows around it. That gap, more than the model's raw ability, seems to decide outcomes in practice.

The zebra results are real. When 20 clinicians worked through 302 difficult NEJM cases, those with LLM access put the correct diagnosis in their top 10 about 52% of the time, compared with 36% for those using search alone. The authors credit the model's wider range of possibilities: it suggests conditions a busy clinician might not think of Does LLM assistance help clinicians build better differentials?. Another study found an LLM beating hundreds of physicians on diagnosis, triage and management, and it included an emergency room study as well as written vignettes Can language models reason better than physicians at diagnosis?. So the ability to recall rare conditions is a real strength. But zebras reward exactly that ability, and routine practice mostly doesn't. Most visits involve common conditions, where the hard part is consistency, follow-through and acting on the advice, not range.

The gap shows up clearly elsewhere. In a UK trial with 1,298 participants, models working alone identified the right condition 94.9% of the time. Members of the public using the same models got it right only 34.5% of the time, no better than people without AI. The knowledge was in the model, but it got lost in how people asked questions and interpreted the answers Why do LLMs fail when users interact with them?. Those participants were laypeople, not clinicians, but the lesson carries over: a benchmark score measures the model, not the model plus the person using it. The strongest real-world evidence points the same way. In nearly 40,000 primary care visits in Nairobi, an LLM safety net cut diagnostic errors by 16% and treatment errors by 13%. The authors credit the design: the tool checked work in the background, the interface fit the clinic's workflow, and the rollout was actively supported. They explicitly say model capability alone does not explain the gains Can AI safety nets reduce errors in live clinical practice?.

There are also reasons to doubt the case scores themselves. Published case challenges sit on the open web, so a model may have seen them during training. In math, one model rebuilt more than half of a well-known benchmark from partial prompts but scored zero on problems published after its training, which shows how memorization can pass for reasoning Does RLVR success on math benchmarks reflect genuine reasoning improvement?. In specialized clinical reasoning tasks, models can also be wrong while sounding confident, and prompting tricks that help in general use don't fix this Why do language models fail confidently in specialized domains?. A useful parallel comes from therapy. LLMs beat trainee therapists on single, isolated responses, but that tells us nothing about ongoing treatment over many sessions Can language models match therapist empathy in real conversations?. Zebra cases have the same limitation: a self-contained puzzle with all the facts provided is very different from a patient whose story comes out over time.

The surprising lesson is that the trait that makes LLMs good at zebras, considering a wide range of possibilities, could be either helpful or harmful in a busy clinic, and the case benchmarks can't tell you which. In Nairobi, that range helped because the tool was built as a quiet safety net rather than a source of answers to sort through. If you want to predict whether a model will help in practice, the corpus suggests looking less at its case-challenge score and more at how it fits into the work: who reads its output, when, and what they do with it.


Sources 7 notes

Does LLM assistance help clinicians build better differentials?

In a study of 20 clinicians on 302 NEJM cases, those with LLM access achieved 51.7% top-10 accuracy versus 36.1% without it. The authors attribute the gain to the LLM's wider differential scope, making lists more comprehensive.

Can language models reason better than physicians at diagnosis?

In physician-adjudicated vignette experiments and an emergency room study, a large language model outperformed hundreds of physicians on differential diagnosis, reasoning, triage, and clinical management tasks across multiple touchpoints.

Why do LLMs fail when users interact with them?

A trial of 1,298 UK participants found GPT-4o, Llama 3, and Command R+ scored 94.9% accuracy alone but users achieved only 34.5% when identifying conditions, no better than controls. The gap lies in how users interpret and act on model suggestions, not model knowledge.

Can AI safety nets reduce errors in live clinical practice?

In 39,849 clinic visits, clinicians with access to an LLM safety net made 16% fewer diagnostic errors and 13% fewer treatment errors than those without. The authors attribute these gains to asynchronous, interface-optimized design and active deployment strategies, not model capability alone.

Does RLVR success on math benchmarks reflect genuine reasoning improvement?

Qwen2.5-Math-7B reconstructs 54.6% of MATH-500 from partial prompts but scores 0.0% on post-release LiveMathBench, revealing dataset contamination. On clean benchmarks, only correct rewards improve performance; random and inverse rewards fail or degrade reasoning ability.

Show all 7 sources
Why do language models fail confidently in specialized domains?

LLMs trained on general text lack sufficient exposure to domain-specific examples, leading to low accuracy paired with high confidence in clinical NLI tasks. Prompting techniques that improved general performance fail to reduce overconfidence in specialized domains.

Can language models match therapist empathy in real conversations?

Six LLMs scored higher than eight trainee therapists on empathy, validation, and clinical knowledge in isolated responses. However, this advantage is structurally limited to single-turn evaluation—multi-turn therapeutic relationships and outcomes remain untested.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.