A prospective clinical feasibility study of a conversational diagnostic AI in an ambulatory primary care clinic

Paper · arXiv 2603.08448 · Published March 9, 2026
Domain Specialization in LLMs

Large language model (LLM)-based AI systems have shown promise for patient-facing diagnostic and management conversations in simulated settings. Translating these systems into clinical practice requires assessment in real-world workflows with rigorous safety oversight. We report a prospective, single-arm feasibility study of an LLM-based conversational AI, the Articulate Medical Intelligence Explorer (AMIE), conducting clinical history taking and presentation of potential diagnoses for patients to discuss with their provider at urgent care appointments at a leading academic medical center. 100 adult patients completed an AMIE text-chat interaction up to 5 days before their appointment. We sought to assess the conversational safety and quality, patient and clinician experience, and clinical reasoning capabilities compared to primary care providers (PCPs). Human safety supervisors monitored all patient-AMIE interactions in real time and did not need to intervene to stop any consultations based on pre-defined criteria. Patients reported high satisfaction and their attitudes towards AI improved after interacting with AMIE (p < 0.001). PCPs found AMIE’s output useful with a positive impact on preparedness. AMIE’s differential diagnosis (DDx) included the final diagnosis, per chart review 8 weeks post-encounter, in 90% of cases, with 75% top-3 accuracy. Blinded assessment of AMIE and PCP DDx and management (Mx) plans suggested similar overall DDx and Mx plan quality, without significant differences for DDx (p = 0.6) and appropriateness and safety of Mx (p = 0.1 and 1.0, respectively). PCPs outperformed AMIE in the practicality (p = 0.003) and cost effectiveness (p = 0.004) of Mx. While further research is needed, this study demonstrates the initial feasibility, safety, and user acceptance of conversational AI in a real-world setting, representing crucial steps towards clinical translation.

Introduction. There is an ever-worsening shortage of primary care physicians (PCPs) affecting virtually every country in the world [1–3]. Tasked with a greater workload and an aging population, burnout rates have surged among PCPs[4]. Technology that facilitates effective use of electronic health records (EHRs) and improves medical team efficiency holds promise for improving accessibility of care and reducing physician burnout[5, 6], while digital care pathways involving telehealth encounters and artificial intelligence (AI) systems for smart intake have become increasingly popular [7, 8]. Large language models (LLMs) show particular promise for extending care availability, with capabilities to engage in nuanced clinical reasoning and conversation [9–11]. Real-world deployments of patient-facing conversational AI further indicate that such systems can meaningfully contribute to patient care coordination and other intake tasks [12, 13], though they have not yet been extensively tested for clinical settings with real human-in-the-loop workflows with PCPs.

Our previous work introduced the Articulate Medical Intelligence Explorer (AMIE), an LLM-based system optimized for clinical dialogue [15–18]. Through studies simulating Objective Structured Clinical Examinations (OSCEs) with trained patient actors, AMIE demonstrated proficiency that was comparable, and in some aspects superior, to human PCPs in diagnostic reasoning during simulations of initial encounters, longitudinal disease management across multiple visits, and encounters requiring clinical reasoning over multimodal artifacts of care [15–17]. Despite these promising results in simulated consultations, the safe and effective translation of such AI systems into real-world clinical practice remains underexplored. The capability to perform a high-quality diagnostic conversation is just one component of safe deployment in care delivery. For AI systems to be effective, they must also adhere to strict safety criteria, especially since emerging evidence suggests that LLM care plans have the potential to cause harm, undertriage medical concerns or unreliably address mental health issues, and that patients are not always able to distinguish between inadequate and adequate medical advice [19–21].

Real-world patient interactions introduce complexities not fully captured by standardized actors or scenarios, including diverse communication styles, unpredictable clinical presentations, a spectrum of health and technology literacy, emotions such as anxiety, and varying levels of management urgency [22]. Collectively, these real-world complexities raise the bar for building conversational AI that remains consistently and reliably safe. Furthermore, assessing the integration of AI tools into existing clinical workflows and gauging the perspectives of both patients and clinicians are vital steps for successful adoption.

To bridge the gap between simulated evaluations and clinical application, we conducted a prospective, single-arm feasibility study of AMIE performing pre-visit clinical conversations within a real-world clinical setting. AMIE collected a detailed history directly from real patients seeking urgent care and presented them information related to possible diagnoses to prepare them for their appointment. This information was then delivered to their PCPs prior to their appointments. Given the high stakes nature of this workflow, our primary objectives were to evaluate the safety of AMIE engaging in conversational history taking with actual patients scheduled for urgent care appointments at a primary care clinic within a leading, high-volume academic medical center, as well as the overall quality of these conversations and operational factors required for successful study completion.

Our contributions are summarized in Figure 1:

• We designed an AMIE system inspired by previous work [14] for this real-world clinical study setting, leveraging the more recent family of Gemini 2.5 base models [23] with Thinking Mode enabled, alongside a refined state-aware chain-of-reasoning strategy that was adapted for robust pre-visit conversational clinical history-taking with patients through iterations of clinical testing and feedback (Figure 1.a).

• Using this adapted version of AMIE, we conducted a pre-registered prospective feasibility study evaluating AMIE interacting with 100 patients in a real-world clinical workflow embedded within an ambulatory primary care clinic. Based on the information shared during the AI encounter, AMIE presented information about potential diagnoses and next steps that their clinician may want to discuss with them. The conversation transcript, along with an automatically generated summary, was provided to the PCP prior to the urgent care appointment. We also developed a safety oversight protocol for all AMIE encounters with prespecified criteria for safety interruptions from supervising physicians (Figure 1.b).

Related work. 5.1. Differential generators Efforts to develop clinical decision support systems have spanned decades, long predating modern machine learning or generative AI systems [33]. The Cornell Medical Index (CMI), developed in 1949 [34], captured information not elicited during physicians’ history-taking that proved pertinent in the diagnostic reasoning process. Ledley & Lusted [35] introduced the principles of leveraging logic, probability, and value theory in medical reasoning to make appropriate diagnostic and management decisions. This conceptual foundation laid the groundwork for subsequent systems like INTERNIST- I, a computerized diagnostic tool developed in the 1970s that could make complex diagnoses in internal medicine, and its successor Quick Medical Reference (QMR) [36, 37]. Diagnostic decision support systems, commonly termed differential generators, such as DxPlain and Isabel performed comparably with expert consensus, though with limited real-world impact on patient care [38–42]. Early evaluations of LLMs suggested performance consistent with differential generators [43–46], and modern reasoning models have far outpaced these systems [10].

5.2. Patient-facing artificial intelligence Patient-facing AI has evolved tremendously over the last decade. Prior to the advent of LLMs, Babylon Health’s Triage and Diagnostic System represented a pioneer in this realm, offering an AI-powered symptom checker that could query and triage patients’ symptoms to provide appropriate next steps [47]. However, evaluation revealed varying diagnostic accuracy, particularly within specialty care A prospective clinical feasibility study of a conversational diagnostic AI in an ambulatory primary care clinic [47–49]. LLMs have drastically advanced the landscape of patient-facing AI, now offering chatbots that can engage in realistic human conversation, gather comprehensive histories, and generate nuanced differentials [16, 17]. Early data from these implementations—largely in telehealth settings—have been promising [48, 50, 51]. Patient communication with AI integrated into health advice lines was shown to be feasible and generally safe [52], and a recent retrospective evaluation by Zeltzer et al. [12] showed that an AI system built from an ensemble of discriminative machine-learning models and augmented with rule-based logic, produced diagnostic and management plans that could potentially meaningfully contribute to patient interactions as an intake tool for telehealth urgent care complaints. LLMs are similarly helpful for care navigation; patients interacting with LLMs for history taking prior to specialist consultation improved efficiency of consultations and perceived care coordination [13].

Method. 2. AI System AMIE is a conversational diagnostic AI system designed to interact with patients in a synchronous text-based interface [17]. The system was built upon Gemini 2.5 Pro (knowledge cutoff January 2025) without any domain-specific fine-tuning. For the purpose of this clinical study, we used the agent setup described in Saab et al. [14] as a starting point, using the Gemini 2.5 family of models [23] with Thinking mode enabled, and further aligned the prompts and conversation phases of this agentic system to adapt it to the specific study setting of AI-driven clinical history-taking prior to an real-world urgent care appointment. This alignment was done based on extensive feedback from clinical experts, as well as synthetic multi-turn roll-outs of dialogues with AI-simulated patients similar to the process described in previous studies [16, 24]. Through this process, we developed an agent which continuously maintains a rich internal state—including an up-to-date patient summary, working differential diagnosis, pertinent information gaps, and draft management plan—to inform its reasoning. Upon meeting a new patient, the system is designed to proactively guide the patient interaction through five distinct phases, with unique prompting and internal logic for each as (Figure 1.a):

• Intake. The agent initiates the consultation, establishing rapport and eliciting basic demographics and the patient’s chief complaint. • History Taking. The agent conducts adaptive inquiry to comprehensively understand the patient’s symptoms and relevant history. Rather than following a static script, questions are dynamically generated based on the agent’s diagnostic hypotheses and information gaps. • Diagnostic Validation. Anchoring on its provisional hypotheses, the agent seeks to fully characterize the patient’s condition, improve its confidence, disambiguate the differential diagnosis, and revise hypotheses as needed, prior to sharing its assessment. Before proceeding, the agent also summarizes its understanding to the patient, inviting corrections or clarifications to ensure data accuracy. • Deliver Assessment. The agent presents possible diagnosis and management options to consider. Given the real-world context of this study, these outputs are framed tentatively—accompanied by clear disclaimers—as possible diagnoses or next steps for the patient to discuss with a provider. • Consultation Wrap-up. The agent confirms the patient’s understanding and provides them space to continue asking questions. It continues to clarify and, in certain cases, update its assessment until the patient elects to conclude the encounter.

After completing a consultation, AMIE generates a conversation summary, which, along with the conversation transcript, was shared with PCPs in our study. For research purposes only, AMIE also generated a management plan, which was not directly shared with the patient or provider, but logged for the purpose of evaluating AMIE.

  1. Methods An overview of our study design including the data collected in the study is provided in Figure 2.

3.1. Study Setting and Oversight We conducted a prospective, single-arm feasibility study to evaluate AMIE’s ability to conduct a real-world pre-visit clinical conversation. The study was performed at Healthcare Associates (HCA), part of Beth Israel Deaconess Medical Center (BIDMC), in Boston, MA from April 2025 to November 2025. HCA is an academic primary care practice with 56 attending physicians, 110 resident physicians, 15 nurse practitioners, and approximately 40,000 total patients.

3.4. Intervention Protocol Following informed consent, patients were scheduled for a remote interaction with AMIE, conducted 0-5 days prior to their planned urgent care visit. Prior to interacting with AMIE, patients completed a brief pre-interaction REDCap survey that captured baseline demographics (age, gender identity, race/ethnicity), health literacy, technology literacy, prior use of chatbots, and the GAAIS (Table 1, B; Figure A.3) [28]. Patients remotely accessed AMIE through a secure synchronous text chat interface.

Discussion. In this work, we performed a feasibility and safety study of a patient-facing conversational AI engaging in urgent care visits in a real world ambulatory primary care setting. We found the deployment of AMIE in a real patient care workflow to be practical with a high patient-AMIE interaction completion rate and follow-up to scheduled urgent care appointments. Additionally, under the supervision of human physicians, we found that conversations with AMIE did not produce any safety alerts based on prespecified study criteria. The quality of AMIE’s communication with patients was highly rated among clinical evaluators along with positive patient sentiment towards AMIE and conversational AI involvement in their care. In this real world setting, we demonstrated PCPs and AMIE had overall similar quality of their differential diagnoses and management plans, with PCPs receiving better ratings on aspects of cost effectiveness and practicality.

6.1. Real World Implementation This study marks the first prospective real-world evaluation of an LLM-based conversational AI agent performing a text-based urgent care visit under real-time supervision of a dedicated safety physician. Prior studies have evaluated conversational AI with real patients but have been either retrospective [12] or limited to advice lines, telehealth, or intake chats prior to a specialist consultation [12, 13, 52]. In contrast, this study included all-comers inclusive of both telehealth and in-office evaluations. Conducting the study in a high-volume academic medical center makes the findings more reflective of real-world implementation within busy clinical workflows. During the study period, approximately 10% of all urgent care visits were enrolled in the study, with patient demographics of study participants skewing towards younger ages, though otherwise largely matching that of the overall urgent care clinic population suggesting generalizability. Regarding the patient-AMIE interaction’s effect on subsequent medical care, only two patients (2%) did not proceed to their scheduled urgent care appointment with a PCP, though there may have been a form of self-selection bias of patients highly engaged in care.

While we found the integration of AMIE to be practical in a busy real world clinical workflow, this was not without challenges. The oversight setup in this study implemented live supervision by a remote physician through screen-sharing of computer screen by the patient via a secure video call. Technical barriers related to this oversight setup were a consistent theme for both patients as well as operationally. Roughly 7% of enrolled patients could not complete the study due to device related issues and upon inquiry of AI supervisors, many patients required significant technology onboarding prior to commencement of the AMIE interaction. These findings are concordant with known health equity barriers of digital health such as low technology literacy and limited access to an adequate device [62, 63]. Our reported rate of interaction incompletion due to patient related technology barriers likely underrepresents the magnitude as our recruitment process deliberately screened out those without an adequate device (laptop or desktop computer) which likely naturally skewed towards a younger patient population. On the operational side, PCPs were able to review the transcript ahead of the scheduled clinic visit only 73% of the time among the cases where PCPs completed the survey. In this study setting, AMIE was not integrated with the clinic’s EHR, but was operated as a separate web application in a secure environment.

Conclusion. In this prospective study, we evaluated the feasibility, safety, and user acceptance of AMIE, a conversational AI system, for conducting clinical history-taking and providing potential diagnoses to patients presenting with urgent concerns in a real-world academic primary care practice setting. In the context of successful safety protocols involving real-time human oversight intended to mitigate risks inherent to introducing novel AI into patient interactions, our findings demonstrate that deployment of AMIE for this task is feasible and conversationally safe, with high rates of successful interaction completion and zero safety stops. The safe presentation of potential diagnoses to patients suggests that conversational AI can meaningfully shift patient-AI interactions from simple information gathering to collaboration and counseling. AMIE also demonstrated strong conversational quality and positive reception from both patients and clinicians. Our results provide initial real-world evidence of AMIE’s clinical reasoning performance.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do clinicians calibrate trust in AI medical recommendations? What prevents LLMs from applying their reasoning knowledge to improve outputs? What limits language model accuracy in evaluating ideas?