When you train an AI further on medical text, does it get better by learning facts or by reasoning through cases?
Does medical fine-tuning help LLMs through knowledge or reasoning ability?
This explores whether training an LLM further on medical material makes it better because it learns more medicine (knowledge) or because it gets better at thinking through cases (reasoning). It also asks whether those two can even be separated.
This explores whether medical fine-tuning helps because the model learns more medicine or because it reasons better through cases. The corpus mostly points to knowledge. In medicine, sharpening reasoning can actually hurt. One line of work finds that LLMs store knowledge mainly in their lower layers and do their reasoning-style adjustment in higher layers. Reasoning training mostly reshapes those upper layers, which may explain why it lifts math scores but can lower scores on knowledge-heavy fields like medicine Why does reasoning training help math but hurt medical tasks?. Medicine rewards recalling the right fact, and math rewards chaining steps, so the same training can push the two in opposite directions.
The tricky part is that fine-tuning isn't very good at adding knowledge either. When models are fine-tuned on facts they didn't already know, they learn those facts slowly. As they do learn them, they start hallucinating more about things they knew before Does fine-tuning on new facts increase hallucination risk?. A broader summary of domain specialization says reinforcement learning mostly prunes inaccurate knowledge rather than adding new capability How do you specialize LLMs without losing general reasoning?. So 'knowledge' here may mean reorganizing and filtering what the model already absorbed in pretraining, not putting new medicine into it. Where the model never saw enough specialist text, the result is low accuracy paired with high confidence on clinical inference tasks, and prompting tricks don't fix it Why do language models fail confidently in specialized domains?.
If you hoped fine-tuning teaches clinical reasoning, the evidence from outside medicine is a warning. Supervised fine-tuning can raise final-answer accuracy while the quality of the reasoning steps falls by about 39%. The answers come out right, but the explanation is built afterward rather than actually leading to the answer Does supervised fine-tuning improve reasoning or just answers?. Fine-tuning on inference tasks can deepen frequency shortcuts instead of teaching real inference Does fine-tuning on NLI teach inference or amplify shortcuts?. RL-trained models still break on slightly altered problems, which suggests template-matching rather than learned procedures Do fine-tuned language models actually learn optimization procedures?. A medical benchmark gain could hide any of these effects.
The deployment studies hint at something you may not have expected. A model's medical value may come less from deep reasoning and more from breadth. Clinicians using an LLM built noticeably better differential diagnoses (51.7% vs 36.1% top-10 accuracy). The authors credit the model for producing wider lists of possibilities, not sharper logic Does LLM assistance help clinicians build better differentials?. Strong models can beat physician baselines on vignettes Can language models reason better than physicians at diagnosis?. Yet when ordinary people used those same models, accuracy fell from 94.9% to 34.5%, no better than controls Why do LLMs fail when users interact with them?. The real bottleneck may be neither knowledge nor reasoning but the handoff to the human.
One caveat: the corpus has only one study that directly separates knowledge from reasoning in medical training, the layer study. Most of the rest is general fine-tuning evidence applied to medicine, so treat the answer 'mostly knowledge, and reasoning training can backfire' as well supported but not settled.
Sources 10 notes
Two-phase inference model shows knowledge retrieval operates in lower network layers while reasoning adjustment happens in higher layers. This separation explains why reasoning training improves math but can degrade knowledge-intensive domains like medicine.
LLMs acquire unknown facts much slower than consistent examples during fine-tuning, but as they master these new facts, they progressively hallucinate more about existing knowledge. This overfitting suggests early-stopping or filtering unknown examples as safer practices.
Research shows supervised fine-tuning raises domain benchmarks but degrades reasoning by 38%, while reinforcement learning prunes inaccurate knowledge rather than adding capability. Every specialization technique has a domain-specific optimal point beyond which performance declines.
LLMs trained on general text lack sufficient exposure to domain-specific examples, leading to low accuracy paired with high confidence in clinical NLI tasks. Prompting techniques that improved general performance fail to reduce overconfidence in specialized domains.
Supervised fine-tuning improves final-answer accuracy on benchmarks but cuts Information Gain by 38.9 percent, meaning models generate correct answers through post-hoc rationalization rather than genuine inferential steps. Standard metrics miss this degradation because they only measure final correctness.
Show all 10 sources
NLI fine-tuning increases LLM reliance on corpus-level frequency patterns (hypernyms more common than hyponyms) rather than semantic relationships. Models perform worse on adversarial cases where frequency patterns contradict actual entailment labels, showing the shortcut was learned more deeply.
Even GRPO-trained models show sharp performance drops on out-of-distribution variants (N-1 test sets) compared to in-distribution problems, indicating RL optimizes template-matching rather than genuine problem-solving procedures.
In a study of 20 clinicians on 302 NEJM cases, those with LLM access achieved 51.7% top-10 accuracy versus 36.1% without it. The authors attribute the gain to the LLM's wider differential scope, making lists more comprehensive.
In physician-adjudicated vignette experiments and an emergency room study, a large language model outperformed hundreds of physicians on differential diagnosis, reasoning, triage, and clinical management tasks across multiple touchpoints.
A trial of 1,298 UK participants found GPT-4o, Llama 3, and Command R+ scored 94.9% accuracy alone but users achieved only 34.5% when identifying conditions, no better than controls. The gap lies in how users interpret and act on model suggestions, not model knowledge.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Towards Accurate Differential Diagnosis with Large Language Models
- Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
- Capabilities of Gemini Models in Medicine
- On the Impact of Fine-Tuning on Chain-of-Thought Reasoning
- Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
- Clinical knowledge in LLMs does not translate to human interactions
- Eliciting Reasoning in Language Models with Cognitive Tools
- Sequential Diagnosis with Language Models