Can conversational AI safely take patient histories without supervision?
A feasibility study tested whether an LLM-based system could conduct real clinical interviews with urgent-care patients without requiring safety interventions. Understanding AI safety in unsupervised clinical settings matters for potential deployment.
A prospective, single-arm feasibility study reports that the Articulate Medical Intelligence Explorer (AMIE), an LLM-based conversational system, can take clinical histories from real patients before urgent-care visits without a single safety intervention. One hundred adults completed a text-chat interaction up to five days before their appointment at an academic primary care practice, between April and November 2025. Human safety supervisors watched every exchange in real time and "did not need to intervene to stop any consultations based on pre-defined criteria." The authors frame this as feasibility and safety rather than effectiveness: the conclusion calls the work "initial real-world evidence" and says further research is needed.
The study also measured more than safety. Patient attitudes toward AI improved after the interaction (p < 0.001). Per chart review eight weeks after the encounter, AMIE's differential diagnosis included the final diagnosis in 90% of cases, with 75% top-3 accuracy. Blinded assessment "suggested similar overall DDx and Mx plan quality" for AMIE and for primary care physicians (PCPs). The gap runs the other way on operations: PCPs outperformed AMIE on the practicality (p = 0.003) and cost effectiveness (p = 0.004) of management plans. The excerpt describes an agent that keeps a running internal state (patient summary, working differential, information gaps, draft plan) and presents its possible diagnoses as "framed tentatively," with "clear disclaimers," for the patient to discuss with a provider. The 90% figure counts whether the final diagnosis appeared anywhere in the list, not whether it ranked first.
Against the library, the result bears most directly on Why do patients distrust medical AI systems?. That note treats the barriers as perceptions, not capabilities, so the blinded parity on differentials bears on the performance barrier, and the attitude shift after a real interaction is direct user-side evidence. The excerpt reports one aggregate attitude result, though, not the three barriers separately, and it says little about accountability beyond the provider-in-the-loop framing. The contrast with Can reinforcement learning personalize which mental health areas to screen? is in what each study counts. CaiTI's therapists flagged GPT-4 for "reading into the user's feelings," a tone-level failure found by clinician review. This study counts stops against pre-specified criteria. Zero stops is a clean count that would not, by itself, reveal that kind of drift.
The excerpt does not establish clinical benefit. There is no control arm and no comparison with usual intake, and the physician comparison rests on written differentials and plans rated blind, not on live visits. The sample is also selected: about 10% of urgent-care visits were enrolled, patients skewed younger, and recruitment screened out anyone without a laptop or desktop. Roughly 7% of enrolled patients could not finish because of device problems, which the authors say likely understates the barrier. PCPs reviewed the transcript before the visit only 73% of the time among those who completed the survey. The 98% who kept their appointments may be inflated by self-selection of engaged patients, as the authors note. The evidence therefore supports feasibility, conversational safety under live physician supervision, and user acceptance at one academic clinic. It does not yet support unsupervised deployment or clinical outcomes, and the practicality and cost gaps mean management-plan parity is not a clean win.
Inquiring lines that read this note 7
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do clinicians calibrate trust in AI medical recommendations?- Does AMIE's advantage hold when patients interact through speech or video instead of text?
- Can an AI system trained on text consultations handle diagnostic uncertainty in real patient encounters?
- What evidence would prove medical AI actually works in clinics?
- What proportion of patients cannot complete AI interviews due to technology barriers?
- Why did primary care physicians review only 73% of AI-generated transcripts?
- Does this colonoscopy finding apply to other medical specialties using AI?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Why do patients distrust medical AI systems?
Explores the psychological barriers that make patients reluctant to adopt medical AI, beyond whether the technology actually works. Understanding these barriers is critical for designing AI systems patients will actually use.
this study's attitude shift and blinded parity bear on the performance and user-side barriers
-
Can reinforcement learning personalize which mental health areas to screen?
Explores whether Q-learning can adaptively prioritize screening across 37 functioning dimensions based on individual patient history, mirroring how therapists naturally focus on areas where clients struggle most.
another clinician-reviewed deployment; safety stops here cannot show the tone-level drift CaiTI's therapists found
-
Can an AI system diagnose better than primary care doctors?
A study compared AMIE, an LLM trained through self-play simulation, against 20 primary care physicians on 149 clinical cases evaluated by specialists. The question asks whether AI can genuinely outperform human doctors in diagnostic reasoning, and what that means for clinical practice.
qualifies: in simulated text chat the same system beat 20 PCPs on differential accuracy, where A's real-clinic blinded review found similar ratings
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- A prospective clinical feasibility study of a conversational diagnostic AI in an ambulatory primary care clinic
- Towards Conversational Diagnostic AI
- AI-based Clinical Decision Support for Primary Care: A Real-World Study
- Clinical knowledge in LLMs does not translate to human interactions
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- Can LLMs identify and repair ruptures? Comparison between clinician practices and LLM behaviors
- People Overtrust AI-Generated Medical Advice despite Low Accuracy
- Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers
Original note title
a conversational AI took histories from 100 real urgent-care patients with zero safety stops — PCPs still won on practicality and cost