Capabilities of Gemini Models in Medicine

Paper · arXiv 2404.18416 · Published April 29, 2024
Domain Specialization in LLMs

Excellence in a wide variety of medical applications poses considerable challenges for AI, requiring advanced reasoning, access to up-to-date medical knowledge and understanding of complex multimodal data. Gemini models, with their strong general capabilities in multimodal and long-context reasoning, offer exciting possibilities in medicine. Building on these core strengths of Gemini 1.0 and Gemini 1.5, we introduce Med-Gemini, a family of highly capable multimodal models that are specialized in medicine with the ability to seamlessly integrate the use of web search, and that can be efficiently tailored to novel modalities using custom encoders. We evaluate Med-Gemini on 14 medical benchmarks spanning text, multimodal and long-context applications, establishing new state-of-the-art (SoTA) performance on 10 of them, and surpass the GPT-4 model family on every benchmark where a direct comparison is viable, often by a wide margin. On the popular MedQA (USMLE) benchmark, our best-performing Med-Gemini model achieves SoTA performance of 91.1% accuracy, using a novel uncertainty-guided search strategy, outperforming our prior best Med-PaLM 2 by 4.6%. Our search-based strategy generalizes with SoTA performance on complex diagnostic challenges from the New England Journal of Medicine (NEJM) and the GeneTuring benchmark. On 7 multimodal benchmarks including NEJM Image Challenges and MMMU (health & medicine), Med-Gemini improves over GPT-4V by an average relative margin of 44.5%. We demonstrate the effectiveness of Med-Gemini’s long-context capabilities through SoTA performance on a needle-in-a-haystack retrieval task from long de-identified health records and medical video question answering, surpassing prior bespoke methods using only in-context learning. Finally, Med-Gemini’s performance suggests real-world utility by surpassing human experts on tasks such as medical text summarization and referral letter generation, alongside demonstrations of promising potential for multimodal medical dialogue, medical research and education.

Introduction. Medicine is a multifaceted endeavor. A clinician’s day-to-day work involves patient consultations, where clear communication of diagnoses, treatment plans, and empathy are essential for building trust. Complex cases necessitate deeper understanding of the patient’s history within the electronic medical record, along with multimodal reasoning from medical images and other diagnostics. To guide their decisions under uncertainty, clinicians must stay abreast of the latest medical information from a wide variety of authoritative sources that can range from research publications to procedural videos. The art of care delivery hinges on a clinician’s ability to perform advanced clinical reasoning, synthesize complex information from diverse and multimodal sources, and collaborate effectively with other clinicians to help people in their care journeys. Although artificial intelligence (AI) systems can assist individual medical tasks (Rajpurkar et al., 2022) and demonstrate early promise towards multimodal multi-task “generalist” medical uses (Moor et al., 2023a; Tu et al., 2024a), the development of more sophisticated reasoning, multimodal, and long-context understanding capabilities would enable significantly more intuitive and helpful assistive tools for clinicians and patients alike.

The advent of large language models (LLMs) and large multimodal models (LMMs), like GPT- 4 (Achiam et al., 2023), PaLM (Chowdhery et al., 2023) and Gemini (Gemini Team, Google, 2023), showed that such models effectively encode clinical knowledge and can perform impressively in medical question answering benchmarks, even for complex cases and scenarios requiring specialized knowledge (Antaki et al., 2023; Eriksen et al., 2023; Kanjee et al., 2023). However, performance on such tasks is far from indicative of real-world utility. The unique nature of medical data and the critical need for safety demand specialized prompting (Nori et al., 2023), fine-tuning, or potentially both along with careful alignment of these models (Ouyang et al., 2022).

Medically fine-tuned LLMs (Luo et al., 2022; Singhal et al., 2023a; Toma et al., 2023) can also provide high-quality long-form answers to nuanced and open-ended medical questions asked by millions of internet users, with Med-PaLM 2 surpassing physicians on axes such as factuality, reasoning, harm, and bias (Singhal et al., 2023b). The potential extends beyond question answering. LMMs (Li et al., 2024; Moor et al., 2023b) such as Flamingo-CXR and Med-PaLM M are comparable with radiologists in controlled settings for generating radiology reports (Huang et al., 2023; Tanno et al., 2024; Tu et al., 2024a). In the more challenging setting of text-based diagnostic consultations with patient actors, the Articulate Medical Intelligence Explorer (AMIE) model outperformed primary care physicians on several evaluation axes for diagnostic dialogue (Tu et al., 2024b).

Despite these promising results, there are considerable opportunities for improvement in performance. LLMs demonstrate suboptimal clinical reasoning under uncertainty, with confabulations and bias remaining key challenges (Omiye et al., 2023; Umapathi et al., 2023). The use of tools and up-to-date medical information (Zakka et al., 2024) to accomplish medical tasks remains a challenge for LLMs, alongside effective collaboration with clinicians (McDuff et al., 2023). Additionally, their ability to handle complex multimodal medical data (for example, integrating images, videos, and de-identified health records over time) is currently limited (Tu et al., 2024a). Although these capabilities are particularly meaningful in medical applications, improvements in performance might be relevant beyond the medical domain. Tasks and benchmarks developed to measure and accelerate the progress of medical LLMs will be broadly impactful.

The Gemini models, as detailed in the Gemini 1.0 and 1.5 technical reports (Gemini Team, Google, 2023, 2024), are a new generation of highly capable multimodal models with novel foundational capabilities that have the potential to address some of these key challenges for medical AI. The models are transformer decoder models (Brown et al., 2020; Vaswani et al., 2017) enhanced with innovations in architecture, optimization and training data, enabling them to exhibit strong capabilities across various modalities including images, audio, video, and text. The recent addition of the mixture-ofexperts architecture (Fedus et al., 2022; Shazeer et al., 2017) allows the Gemini models to efficiently scale and reason over significantly longer and more complex data at inference time.

Building on the strengths of the Gemini models, we present Med-Gemini, a family of models fine-tuned and specialized for medicine. The notion of generalist medical AI models has received considerable attention with impressive demonstrations of the possibilities for such systems (Tu et al., 2024a).

Method. As introduced in the Gemini technical reports (Gemini Team, Google, 2023, 2024), the Gemini ecosystem encompasses a suite of models varying in size, modality encoders, and architectures, trained on a wide variety of high quality data across many modalities. The Gemini models exhibit state-of-the-art results across a diverse array of language, reasoning, coding, multilingual, image, and video benchmarks. Notably, the Gemini 1.0 Ultra model excels in language-based tasks that require complex reasoning, and the Gemini 1.5 Pro model adds the ability to efficiently handle and make use of long-context inputs spanning millions of tokens and/or multimodal inputs such as hours of video or tens of hours of audio. Gemini 1.0 Nano is the smallest model variant in the Gemini model family that can run efficiently on-device.

We develop our Med-Gemini models by building on the Gemini family, focusing on the following capabilities and methods:

  1. Advanced reasoning via self-training and web search integration: For language tasks that require less complex reasoning, such as summarizing medical notes and creating referral letters, Capabilities of Gemini Models in Medicine we introduce Med-Gemini-M 1.0 by fine-tuning the Gemini 1.0 Pro model. For other tasks that require more advanced reasoning, we introduce Med-Gemini-L 1.0 by fine-tuning the Gemini 1.0 Ultra model using a self-training method to enable the models to efficiently use web search. We develop a novel uncertainty-guided search strategy at inference time to improve performance on complex clinical reasoning tasks. 2. Multimodal understanding via fine-tuning and customized encoders: The Gemini models are natively multimodal and have demonstrated impressive zero-shot performance on many multimodal benchmarks. However, the unique nature and heterogeneity of some medical modalities require fine-tuning to achieve the best possible performance. We introduce Med- Gemini-M 1.5 by performing fine-tuning with Gemini 1.5 Pro on a suite of multimodal medical datasets. We introduce Med-Gemini-S 1.0 and demonstrate the Gemini models’ capability to adapt to novel medical modalities using specialized encoders with the Gemini 1.0 Nano model. 3. Long-context processing with chain-of-reasoning: For the long-context processing tasks, we re-use Med-Gemini-M 1.5 with a long-context configuration. In addition, we also develop a novel inference-time chain-of-reasoning technique inspired by Tu et al. (2024b) to enable better understanding of long EHRs.

2.1. Advanced reasoning via self-training and web search integration Clinical reasoning is a fundamental skill that underpins successful care. Although it is a broad field with many definitions, clinical reasoning can be conceptualized as an iterative process by which a physician integrates their own clinical knowledge with initial patient information to form a case representation. This representation is then used to guide the iterative acquisition of additional information until a confidence threshold is reached to support a final diagnosis with plans for treatment and management (Gruppen, 2017). During this process, a physician may reason across many diverse inputs, such as patient symptoms, medical and socio-economic history, investigations and lab tests, prior responses to treatments and other wider factors such as epidemiological data. Moreover, many of these inputs have a time component, such as a series of evolving symptoms, lab measurements over time, or the various temporal data that is collected for monitoring health, such as electrocardiograms (ECGs). Medical knowledge is highly non-stationary, with reducing “doubling times” in the volume of medical information driven by the rapid pace of research (Densen, 2011; Grandage et al., 2002). To ensure that their outputs reflect the latest information in this domain, LLMs might ideally not only possess strong reasoning capabilities but also be able to integrate up-to-date information, for example, from authoritative web sources. This grounding in external knowledge has the potential to reduce uncertainty in the model’s responses, but requires an informed approach to information retrieval itself.

Discussion. Med-Gemini, built upon the Gemini models, demonstrates significant advancements in clinical reasoning, multimodal understanding, and long-context processing within the medical domain. This is evidenced by its strong performance across a diverse range of 25 tasks spanning 14 medical benchmarks, encompassing medical knowledge, clinical reasoning, genomics, waveforms, medical imaging, health records and videos.

MedQA performance Notably, Med-Gemini-L 1.0 achieves a new SoTA on MedQA (USMLE), a popular benchmark for medical question answering with the use of self-training based fine-tuning and search integration. Our thorough relabeling of the MedQA test set (performed by attending clinicians) reveals important insights. While MedQA (USMLE) is a useful benchmark for assessing medical knowledge and reasoning, it is essential to acknowledge its limitations. We discover that approximately 4% of the questions contain missing information, and an additional 3% potentially have labeling errors. Establishing definitive ground truth is frequently challenging in medicine, where inter-reader variability and ambiguity are common and medical knowledge is constantly evolving. Our observations suggest that further improvements in SoTA performance on the MedQA (USMLE) benchmark in isolation may not directly correlate to progress in the capabilities of medical LLMs for meaningful real-world tasks and as such it is important to perform more comprehensive benchmarking and evaluation representative of real-world clinical workflows (Fleming et al., 2023). In general, most benchmarks have limitations around dataset size and quality. While we focus our analysis here on MedQA (USMLE), prior work has suggested similar issues with other popular benchmark datasets (Xu et al., 2023). Retraining Med-Gemini-M 1.5 with a new split of the PAD-UFES-20 dermatology dataset leads to a drop of 7.1% as compared to our results in Table 2. As such, careful attention needs to be given to the size and quality of datasets when interpreting and contextualizing model performance.

Web search integration Med-Gemini’s integration with web search presents exciting possibilities to provide more factually accurate and reliable answers to medical queries with LLMs. In this work, we focus on training Med-Gemini-L 1.0 to issue web search queries when uncertain and integrate the results when producing responses. While the results on MedQA, NEJM CPC, and GeneTuring benchmarks are promising, significant further research is necessary. For example, we haven’t considered restricting the search results to more authoritative medical sources (Zakka et al., 2024), using multimodal search retrieval or performed analysis on accuracy and relevance of search results and the quality of the citations (Wu et al., 2024). Further, it remains to be seen if smaller LLMs can also be taught to make use of web search. We leave these explorations to future work.

Promising multimodal conversational capabilities The multimodal conversational capabilities of Med-Gemini-M 1.5 are promising given they are attained without any specific medical dialogue fine-tuning. Such capabilities allow for seamless and natural interactions between people, clinicians, and AI systems. As showcased in our qualitative examples, Med-Gemini-M 1.5 has the capability to engage in multi-turn clinical dialogues, request additional information such as images when needed, explain their reasoning in a comprehensible manner, and even help provide information useful for clinical decisions while appropriately deferring the final decision to human experts. This capability has significant potential for helpful real-world applications, including assisting clinicians and patients, but of course also entails highly significant associated risks.

Conclusion. Large multimodal language models are ushering in a new era of possibilities for health and medicine. The capabilities demonstrated by Gemini and Med-Gemini suggest a significant leap forward in the depth and breadth of opportunities to accelerate biomedical discoveries and assist in healthcare delivery and experiences. However, it is paramount that advancements in model capabilities are accompanied by meticulous attention to the reliability and safety of these systems. By prioritizing both aspects, we can responsibly envision a future where the capabilities of AI systems are meaningful and safe accelerators of both scientific progress and care in medicine.

Limitations. Responsible AI Our work has been primarily focused on capabilities and improvements and the art of the possible with Gemini models. An important focal area for future exploration is the integration of the responsible AI principles throughout the model development process (Pfohl et al., 2024), including, but not limited to, the principles of fairness, privacy, equity, transparency and accountability. Privacy considerations in particular need to be rooted in existing healthcare policies and regulations governing and safeguarding patient information. Fairness is another area that may require attention, as there is a risk that AI systems in healthcare may unintentionally reflect or amplify historical biases and inequities (Abràmoff et al., 2023; Char et al., 2018; Cirillo et al., 2020; Gichoya et al., 2022; Obermeyer et al., 2019; Pfohl et al., 2024), potentially leading to disparate model performance and harmful outcomes for marginalised groups. Such health disparities have been identified across gender (Kent et al., 2012), race (Obermeyer et al., 2019; Williams and Wyatt, 2015), ethnicity (Razai et al., 2021), socioeconomic status (Steptoe and Zaninotto, 2020), sexual orientation (Medina-Martínez et al., 2021), age (Jackson et al., 2019), and other sensitive and/or protected personal characteristics.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do clinicians calibrate trust in AI medical recommendations? What prevents LLMs from applying their reasoning knowledge to improve outputs? How do curriculum design and feedback approaches affect model learning? Can language models reliably simulate personas and predict behavior? Why do language models fail at sustained therapeutic relationships despite understanding techniques? What limits language model accuracy in evaluating ideas? Can confidence signals reliably detect flawed reasoning in language models? Does scaling reasoning capability create fundamental tradeoffs in control and reliability? Can external verification systems adequately replace learned reasoning in AI outputs? What explains the gap between benchmark scores and true reasoning capability? Can smaller specialized models match frontier models on key metrics?