People Overtrust AI-Generated Medical Advice despite Low Accuracy
Source: Shekar, Pataranutaporn, Sarabu, Cecchi, Maes (MIT Media Lab) — NEJM AI · 2025-05-22
BACKGROUND This article presents a comprehensive analysis of how artificial intelligence (AI)–generated medical responses are perceived and evaluated by nonexperts.
METHODS We conducted a study in which a total of 300 participants gave evaluations for medical responses that were either written by a medical doctor on an online health care plat form or generated by a large language model and labeled by physicians as having high accu racy or low accuracy.
RESULTS Results showed that participants could not effectively distinguish between AI-generated responses and doctors’ responses and demonstrated a preference for AI-generated responses, rating high-accuracy AI-generated responses as significantly more valid, trustworthy, and complete/satisfactory. Low-accuracy AI-generated responses on average performed very similarly to doctors’ responses. Participants not only found these low-accuracy AI-generated responses to be valid, trustworthy, and complete/satisfactory, but also indicated a high tendency to follow the potentially harmful medical advice and incorrectly seek unnecessary medical attention as a result of the response provided. This problematic reaction was comparable with, if not stronger than, the reaction they displayed toward doctors’ responses. Both experts and nonexperts exhibited bias, finding AI-generated responses to be more thorough and accurate than doctors’ responses but still valuing the involvement of a doctor in the delivery of their medical advice.
CONCLUSIONS The increased trust placed in inaccurate or inappropriate AI-generated medical advice can lead to misdiagnosis and harmful consequences for individuals seek ing help. Further, participants were more trusting of high-accuracy AI-generated responses when told they were given by a doctor, and experts rated AI-generated responses signifi cantly higher when the source of the response was unknown. Ultimately, AI systems should be implemented in collaboration with medical professionals when used for the delivery of medical advice in order to prevent misinformation while reaping the benefits of such cut ting-edge technology.
Introduction. University of California San Diego Health; University of Wisconsin Health in Madison, Wisconsin; and Stanford Health Care were among the first organizations to deploy technology to respond to health care messages automatically.42,43 Studies have shown notable performances of LLMs com pleting medical licensing exams.44-49 One study showed that GPT-4 exceeds the passing score of the official prac tice materials for the United States Medical Licensing Examination.44 Another study found that ChatGPT was able to generate higher quality and more empathetic responses to patient questions.50 A randomized controlled trial for medical diagnosis comparing physicians alone, AI alone, and physicians augmented with AI had the unex pected finding that AI alone outperformed all the other groups.51 However, a follow-up study from the same group found that physicians augmented with AI performed com parably to AI alone and both groups together outperformed physicians not using AI.52 Despite their potential benefit to health care and medi cine,53-55 the stochastic nature of LLMs makes it challeng ing to determine when LLMs will give factually correct answers versus confidently provide false information (i.e., hallucination or confabulation).56,57 The stakes are high in medical applications. For instance, a study on the use of LLMs to select next-step antidepressant treatment in major depression showed that the model’s inclusion of less opti mal clinical recommendations posed a significant risk if used routinely without expert supervision.23 As LLMs become more prevalent in mainstream search engines and conversational interfaces, it is not always fea sible to have expert supervision. Simply focusing on the accuracy of LLMs in answering medical questions is insuf ficient, as this fails to capture the broader implications of the technology for the health care system and society at large.53,58,59 We argue that it is critical to study how the lay public perceives, evaluates, and is affected by AI-generated responses, especially when incorrect, as LLM nonex perts will encounter situations where they might trust AI-generated advice, particularly in the absence of immedi ate medical professional guidance. Overrelying on false or incomplete AI-generated responses could lead to delayed or inappropriate treatment, potentially worsening health outcomes and even endangering lives.
Related work. University of California San Diego Health; University of Wisconsin Health in Madison, Wisconsin; and Stanford Health Care were among the first organizations to deploy technology to respond to health care messages automatically.42,43 Studies have shown notable performances of LLMs com pleting medical licensing exams.44-49 One study showed that GPT-4 exceeds the passing score of the official prac tice materials for the United States Medical Licensing Examination.44 Another study found that ChatGPT was able to generate higher quality and more empathetic responses to patient questions.50 A randomized controlled trial for medical diagnosis comparing physicians alone, AI alone, and physicians augmented with AI had the unex pected finding that AI alone outperformed all the other groups.51 However, a follow-up study from the same group found that physicians augmented with AI performed com parably to AI alone and both groups together outperformed physicians not using AI.52 Despite their potential benefit to health care and medi cine,53-55 the stochastic nature of LLMs makes it challeng ing to determine when LLMs will give factually correct answers versus confidently provide false information (i.e., hallucination or confabulation).56,57 The stakes are high in medical applications. For instance, a study on the use of LLMs to select next-step antidepressant treatment in major As LLMs become more prevalent in mainstream search engines and conversational interfaces, it is not always fea sible to have expert supervision.
Method. This article presents three experiments investigating AI-generated medical responses to medical questions and whether the AI-generated responses are comparable to phy sician responses. Additionally, this study explores the per ception of these AI-generated medical responses from the perspective of both the public and physicians.
TASK DESCRIPTION First, we investigated whether participants would be able to distinguish AI-generated responses from doctors’ responses as a preliminary understanding of participant perception of AI and doctors in responding to health inqui ries. In this first experiment, 100 online participants were presented with 10 medical question–response pairs ran domly selected from a collection of 30 doctors’ responses, 30 high-accuracy AI-generated responses, and 30 low-ac curacy AI-generated responses. After reading the provided medical question–response pair, participants provided Likert scale evaluations on a scale of 1 (strongly disagree) to 5 (strongly agree) on their understanding of the med ical question and their understanding of the response. Additionally, they indicated their belief about the response source (response given by a doctor or an AI text generator) and provided a Likert scale evaluation of their confidence in the source they selected on a scale of 1 (low confidence) to 5 (high confidence). The full set of questionnaires is listed in the Supplementary Appendix.
For the second experiment, we assessed how participants evaluate responses generated by the AI system compared with those provided by doctors, when they are unaware of the exact source of the responses. This experiment, similarly to experiment 1, involved 100 participants, who were presented with 10 medical question–response pairs randomly selected from a collection of doctors’ responses, high-accuracy AI-generated responses, and low-accuracy AI-generated responses. Here, participants provided Likert scale evaluations on a scale of 1 (strongly disagree) to 5 (strongly agree) on their understanding of the medical ques tion and their understanding of the response. Additionally, they were asked to indicate their perception of response validity (yes/no). Finally, participants provided Likert scale evaluations of the trustworthiness of the response; the completeness and satisfaction of the response; partic ipant tendency to search for additional information based on the response; participant tendency to follow the advice provided in the response; and participant tendency to seek subsequent medical attention as a result of the response.
In the third experiment, we investigated if participants exhibited biases toward or against certain response types. Similarly to experiments 1 and 2, 100 participants were presented with 10 medical question–response pairs ran domly selected from a collection of doctors’ responses, high-accuracy AI-generated responses, and low-accuracy AI-generated responses. However, at the start of the sur vey, participants were randomly shown one of three labels: “The responses to each medical question were given by a %(doctor)”; “The responses to each medical question were given by %(artificial intelligence (AI))”; or “The responses to each medical question were given by a %(doctor assisted by AI).” Participants then provided Likert scale evaluations on a scale of 1 (strongly disagree) to 5 (strongly agree) on their understanding of the medical question and their understanding of the response. They indicated their per ception of response validity (yes/no). Finally, they provided Likert scale evaluations on a scale of 1 (strongly disagree) to 5 (strongly agree) of the trustworthiness of the response; the completeness and satisfaction of the response; partic ipant tendency to search for additional information based on the response; participant tendency to follow the advice provided in the response; and participant tendency to seek subsequent medical attention as a result of the response.
Discussion. Participants demonstrated similar understanding of med ical questions across all groups (high-accuracy AI, low-ac curacy AI, and doctor). This consistency, combined with no differences in linguistic characteristics between response types, thus controlling for any confounding factors related to medical question–response linguistics that could impact evaluation outcomes, ensures that evaluation differ ences reflect perception of responses rather than question comprehension.
Participants displayed an approximate 50% accuracy rate in discerning the origin of the medical responses, making it clear that they struggled to effectively differentiate between medical advice offered by a doctor and medical responses generated by AI. This holds true even when the accuracy of the AI-generated medical response is comparatively low. Thus, participants perceived the AI-generated responses as remarkably similar to those provided by doctors, rendering them unable to accurately differentiate between the advice given by the AI and that offered by a registered physician on the online health care platform HealthTap.
In addition to participants’ inability to distinguish AI-generated responses from doctors’ responses, we found that participants evaluated AI-generated responses as almost equal to, if not better than, responses pro vided by doctors across all metrics. AI-generated medical responses were found to be as comprehensive, valid, trust worthy, complete/satisfactory, and persuasive as doctors’ responses, with AI-generated responses of high accuracy performing significantly better in a majority of the met rics. Furthermore, on average, albeit not significantly, low-accuracy AI-generated responses presented a higher level of performance than the doctors’ responses across all the evaluation metrics.
Participants’ inability to differentiate between the quality of AI-generated responses and doctors’ responses, regardless of accuracy, combined with their high evaluation of low-ac curacy AI responses, which were deemed comparable with, if not superior to, doctors’ responses, presents a concerning threat. When unaware of the response’s source, participants are willing to trust, be satisfied, and even act upon advice provided in AI-generated responses, similarly to how they would respond to advice given by a doctor, even when the AI-generated response includes inaccurate information. This unexpected trust and satisfaction with low-accuracy AI-generated responses may lead to unwitting acceptance of harmful or ineffective medical advice and concerns of liability for any resulting adverse patient outcomes.63 Participants evaluating unlabeled medical responses favored AI-generated ones, trusting even low-accuracy responses. However, source labeling changed evaluations significantly. High-accuracy AI responses labeled as doctor were deemed more trustworthy than when labeled as AI, suggesting that, while participants appreciate AI-generated advice, they gen erally prefer receiving it from doctors. Notably, the doctor label alone did not enhance perception of low-accuracy AI responses. This effect was strongest with high-accuracy AI responses, demonstrating a combined effect of desirable source and high-accuracy model for achieving desirable evaluations. Similar patterns appear in other domains, as shown in a study where humanlike explanations combined with high-accuracy AI responses increased trust in a legal decision-making advice.64 Interestingly, our expert evalu ators showed similar bias, rating AI responses significantly higher when source-blind. This reveals that even those responsible for establishing objective truth and assessing model efficacy can be susceptible to inherent biases.
Conclusion. Our study reveals that participants rate AI-generated responses, particularly high-accuracy ones, as equal to or better than doctors’ responses across all metrics, while maintaining higher trust in responses attributed to doc tors. However, responses labeled as doctor assisted by AI showed no significant improvement in evaluations, com plicating the ideal solution of combining AI’s comprehen sive responses with physician trust. This underscores the complexities of the situation and emphasizes the intricate dynamics through which participants and experts inter act and perceive medical responses. Future doctor-assist ed-by-AI applications will need careful framework design to build trust.58 Our findings suggest three key considerations: AI can effectively deliver medical responses when accurate; inaccurate AI responses risk misleading the public through persuasive humanlike language; and expert oversight is crucial to maximize AI’s unique capabilities while minimiz ing risks. Health care providers should focus on developing best practices for AI integration, establishing transparent physician involvement in AI-generated information deliv ery, and creating standardized methods for evaluating doc tor–AI collaboration across medical domains.
This research reveals that AI-generated medical responses are not only indistinguishable from doctors’ responses but are often preferred by the general public across all metrics — understandability, validity, trust, completeness/satisfac tion, and persuasion. However, participants showed higher trust when they believed responses came from doctors. This creates a concerning paradox: while AI responses can be compelling and seemingly trustworthy, their potential inaccuracies could lead to harmful or fatal consequences if used without expert oversight. These findings suggest that integrating AI into medical information delivery requires a more nuanced approach than previously considered.
Limitations. There are several limitations to this study. First, this study uses GPT-3 rather than a more recent version of the model. While newer models might improve accuracy, it is concern ing that even low-accuracy responses from older models proved convincing.
Second, our participant pool, recruited through an online platform, may be skewed toward the technologically savvy and represents mainly those 18 to 49 years of age. Additionally, participants evaluated hypothetical scenarios rather than their own medical questions, lacking personal investment in the responses. Third, the study examines single question–response pairs without the context and follow-up typical in real clinical scenarios, where doctors would likely request additional information before provid ing advice. Future research should explore how such con text affects AI’s role in medical question answering.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How do clinicians calibrate trust in AI medical recommendations?- Why did lay users with AI models fail to match unaided physicians on diagnosis?
- Do expert physicians also prefer AI-written medical text when it is unlabeled?
- Why does labeling advice as AI from a doctor change how people trust it?
- Can people tell which medical advice is accurate based only on how it reads?
- How does trusting wrong AI advice change what medical action people decide to take?
- Do people judge AI moral advice differently than human advice when they know the source?
- Does positive sentiment bias in AI content harm information quality?
- Can cognitive governance help users interpret AI outputs better?
- Why do people prefer AI moral arguments when they don't know the source?
- Why do users override their own judgment when AI says a headline is false?
- How does AI presentation authority substitute for actual expert judgment?
- Why do users default to treating AI outputs as equally reliable evidence?
- What happens to expert credibility when AI-generated claims drown out specialist signals?
- Does surface authority without earned authority create risks in expert judgment?
- Can humans develop oversight strategies that work across all GenAI rhetorical shifts?
- What assumptions about oversight fail when AI acts as rhetorical interlocutor?
- How does AI reliance change professional judgment and autonomy?
- Does AI assistance actually reduce neural processing and brain connectivity over time?
- How does incremental AI use gradually reduce human decision-making capacity?
- Why do users believe they produced independent competence when they actually used AI assistance?