SYNTHESIS NOTE
Topics›Domain Specialization›this note

Does medical pretraining actually improve model performance?

When medical models are fairly compared to their base counterparts with per-model prompt tuning and statistical significance testing, do they consistently outperform on medical question answering tasks?

Synthesis note · 2026-10-06 · sourced from Domain Specialization

The paper argues that medical domain-adaptive pretraining (DAPT), the common practice of continuing to pretrain an open general-domain model on PubMed and medical textbooks, does not reliably make models better at medical question answering. Comparing seven medical LLMs and two medical VLMs against their own base models, it finds that "all medical VLMs and nearly all medical LLMs fail to consistently improve over their general-domain counterparts" in zero- and few-shot prompting. The headline figure is for the 3-shot setting: across tasks and model pairs, medical LLMs beat their base models in 12.1% of cases, tie in 49.8%, and are significantly worse in the remaining 38.2%. The paper's closing claim is the stronger reading: "state-of-the-art general-domain models may already exhibit strong medical knowledge and reasoning capabilities."

The mechanism is comparison design, not a new benchmark. The paper makes three moves: it compares each medical model head-to-head with the base model it was derived from, so that only the medical pretraining differs; it selects prompt format and few-shot examples separately for each model on the validation set; and it reports bootstrapped confidence intervals on the test set. The second move matters because the paper notes that the "optimal" prompt choice "rarely correlates between different models," so a single prompt tuned for the medical model can flatter it. The zero-shot numbers show the size of the effect. When the prompt is optimized only for the medical model and uncertainty is ignored, medical LLMs and VLMs outperform their base models on average in 70.5% and 62.5% of QA tasks. After per-model prompt selection and significance testing, the improvements are significant in only 9.4% and 6.3% of tasks. The paper says its ablations show these practices "substantially impact conclusions."

Against the nearby notes, this is a measured test of one route in the knowledge-injection taxonomy. How do knowledge injection methods trade off flexibility and cost? places domain expertise embedded in weights during pretraining or fine-tuning on the static side, where the cost is training and the payoff is inference speed. Continued pretraining on medical text is that route, and this paper finds its payoff over the base model is largely absent in closed-ended QA once comparisons are fair. It also qualifies the knowledge-dominance reading in Does medical AI need knowledge or reasoning more?. Medicine may be knowledge-dominant, but if general models already hold that knowledge, the gap may lie in eliciting it rather than injecting it. The two readings are compatible: the paper's own phrasing is that general models can be "leveraged effectively when prompted appropriately."

The excerpt does not establish more than this. It covers closed-ended QA scored by exact match under greedy decoding. The authors concede that open-ended consumer health QA and information extraction from clinical notes are untested, and that "it is certainly possible that an analysis like ours would find improved performance on such tasks." The model set is not exhaustive, and the excerpt does not show the per-pair results behind the BioMistral-7B exception. The implication, at the strength the evidence allows: on closed-ended medical QA, a medical checkpoint should not be credited over its own base model until both are compared under per-model prompts with uncertainty reported. The result does not show that DAPT is useless for clinical extraction or for open-ended tasks.

Inquiring lines that read this note 5

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do curriculum design and feedback approaches affect model learning? How do clinicians calibrate trust in AI medical recommendations? How does fine-tuning trade off accuracy against reasoning quality?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 87 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

medical domain-adaptive pretraining largely disappears once each model gets its own prompt and statistical uncertainty is counted