Strip the scans and lab tables from a patient case and leave only text: how much of an AI's diagnostic edge survives?
How much does missing images and tables limit LLM diagnostic reasoning?
This explores how much an LLM's diagnostic reasoning suffers when a case reaches it as text only, without the X-rays, scans, lab tables and charts a clinician would normally see, and whether the strong results reported for medical LLMs hold up once that missing material is taken into account.
This explores how much an LLM's diagnostic reasoning suffers when a case reaches it as text only, without the images and tables a clinician would normally look at. The collection doesn't directly measure this. None of these notes compares the same cases with and without the visual material. What it does have is enough on either side of the gap to show why the question matters, and why the headline results should be read with it in mind.
The headline results are strong. Against a baseline of hundreds of physicians, one LLM outperformed doctors on differential diagnosis, triage and management, in written vignettes and in an emergency-room study Can language models reason better than physicians at diagnosis?. In a separate study of 302 NEJM cases, clinicians with LLM access got the right diagnosis into their top 10 51.7% of the time, versus 36.1% with search alone. The authors credit the LLM's wider list of possibilities Does LLM assistance help clinicians build better differentials?. Notice what that explanation says: the gain comes from breadth of recall over text, not from reading an image well. These are mostly studies of reasoning from a written description of a case. That is a real skill, but it isn't the whole job.
The lateral evidence suggests that missing visuals may cost more than a few percentage points. One note finds that in multimodal models, the bottleneck in fine-grained visual tasks is where the model directs its visual attention, not how well it explains itself. Long chain-of-thought reasoning, the technique that helps with text problems, actually makes perception worse Does verbose chain-of-thought actually help multimodal perception tasks?. So adding images doesn't simply plug into the reasoning ability those diagnostic studies show. Perception is a separate skill, and training aimed at better reasoning can work against it.
The more troubling issue is what the model does when information is absent. LLMs often fail not because they lack knowledge but because they don't bring up unstated conditions that matter. Prompts that force them to list those conditions explicitly raised accuracy from 30% to 85% Do language models fail at identifying unstated preconditions?. A missing CT scan is exactly this kind of unstated condition: a careful clinician would say "I can't rule this out without the imaging," and a model may never mention it. In specialized clinical tasks, models also tend to be wrong while sounding confident, and prompting tricks that improve general accuracy don't reduce that overconfidence Why do language models fail confidently in specialized domains?. Put these together and the risk is not a model that refuses to answer. It's a model that fills the gap with a fluent, plausible diagnosis. The fabrication framing explains why: correct and incorrect outputs come from the same text-generating process, so nothing internal signals that a key input was missing Should we call LLM errors hallucinations or fabrications?.
The takeaway: the useful question may be less about how much accuracy you lose without the image, and more about whether the model notices that the image is missing. The evidence suggests it often won't unless you ask it to. A practical safeguard is to prompt the model to list what information it is reasoning without before it commits to a diagnosis. For a direct measurement of text-only versus full-case performance, the collection would need papers it doesn't have yet.
Sources 6 notes
In physician-adjudicated vignette experiments and an emergency room study, a large language model outperformed hundreds of physicians on differential diagnosis, reasoning, triage, and clinical management tasks across multiple touchpoints.
In a study of 20 clinicians on 302 NEJM cases, those with LLM access achieved 51.7% top-10 accuracy versus 36.1% without it. The authors attribute the gain to the LLM's wider differential scope, making lists more comprehensive.
Long rationales and text-token RL help reasoning but hurt fine-grained perception tasks because the actual bottleneck is visual attention allocation, not verbalization. Standard CoT optimization trains the wrong policy target.
LLMs struggle not from lacking world knowledge but from failing to bring background conditions forward as relevant constraints. Prompting that forces explicit enumeration of preconditions raises accuracy from 30% to 85%, revealing the frame problem persists in statistical systems.
LLMs trained on general text lack sufficient exposure to domain-specific examples, leading to low accuracy paired with high confidence in clinical NLI tasks. Prompting techniques that improved general performance fail to reduce overconfidence in specialized domains.
Show all 6 sources
LLMs generate text through statistical token relationships without grounding in shared context. Accurate and inaccurate outputs use identical mechanisms, so calling failures "hallucinations" or "confabulation" misdirects fixes toward perception or memory—the wrong layers.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Towards Accurate Differential Diagnosis with Large Language Models
- Capabilities of Gemini Models in Medicine
- Sequential Diagnosis with Language Models
- Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
- Superhuman performance of a large language model on the reasoning tasks of a physician
- The Model Says Walk: How Surface Heuristics Override Implicit Constraints in LLM Reasoning
- Diagnostic Reasoning Prompts Reveal the Potential for Large Language Model Interpretability in Medicine
- Clinical knowledge in LLMs does not translate to human interactions