INQUIRING LINE

AMIE beat primary care doctors in typed chats, but no study yet checks whether that edge survives speech or video.

Does AMIE's advantage hold when patients interact through speech or video instead of text?

This explores whether AMIE, the diagnostic AI that beat primary care doctors in typed consultations, would still come out ahead if patients spoke to it or saw it on video. The collection doesn't test that directly, but it has evidence on why the medium might matter.


This explores whether AMIE's edge over primary care doctors survives a change of medium, from typed chat to voice or video. The short answer: the corpus has no study that runs AMIE through speech or video, so nobody here has tested it. What the corpus does have is a set of clues about where AMIE's advantage comes from and how much the communication channel can change outcomes. Together they suggest the advantage might not carry over intact.

Start with where the advantage came from. In the original study, AMIE beat primary care physicians on 28 of 32 specialist-rated measures across 149 simulated cases, all conducted by text (Can an AI system diagnose better than primary care doctors?). The win came from reasoning over the information it had gathered, not from drawing out the patient's history. That distinction matters for the voice and video question. A typed transcript is AMIE's natural environment, and it may also hold doctors back, because they lose the tone, hesitation and visible distress they normally use to steer an interview. If the gap partly reflects doctors working without their usual tools, adding voice or video could close it from the physician side.

The therapy research shows how much the medium can matter. In a 15-day study, a robot and a paper worksheet both reduced students' distress, while a chatbot running the same language model did not (Why do robots outperform chatbots in therapy despite identical language models?). A broader synthesis says the active ingredient in therapeutic AI is often the sense of a present, non-judging listener rather than clinical technique (Is conversational presence more therapeutic than clinical technique?). Diagnosis is not therapy, but the lesson carries: you can't assume a result measured in text will hold in another format. Fine-grained speech details also carry information. Patient pauses and filler words signal relaxed communication and a stronger working relationship (Does therapist self-reference language predict weaker therapeutic alliance?). A voice-based AMIE would have to read those signals, which typed text strips out.

The closest the corpus gets to real-world deployment is a follow-up in which AMIE took histories from 100 real urgent-care patients with no safety stops. Patients' attitudes toward AI also improved afterward (Can conversational AI safely take patient histories without supervision?). But its management plans fell behind physicians' on practicality and cost. So when AMIE leaves the controlled text setup, its advantage already looks narrower. A related warning comes from therapy research: language models beat trainee therapists on single isolated responses, but that edge hasn't been shown to last over an ongoing relationship (Can language models match therapist empathy in real conversations?). Gains measured in a tidy, limited format often shrink when the setting gets messier.

The takeaway you might not expect: the right question may be less whether AMIE can handle voice and more whether text was helping AMIE by handicapping the doctors. The corpus can't settle that. A head-to-head trial across text, voice and video is the missing experiment.


Sources 6 notes

Can an AI system diagnose better than primary care doctors?

An LLM-based diagnostic system called AMIE exceeded primary care physician performance in text-based simulated consultations across 149 case scenarios, scoring higher on 28 of 32 specialist-rated dimensions. The advantage lay in inference from gathered information rather than in eliciting history.

Why do robots outperform chatbots in therapy despite identical language models?

A 15-day study with 38 students found that robots and worksheets significantly reduced psychological distress while a chatbot using the same LLM did not. The active ingredient was the medium—social presence and structured format—not language capability.

Is conversational presence more therapeutic than clinical technique?

ELIZA matches modern chatbots on symptom reduction, RLHF training degrades emotional attunement, and embodied robots outperform text-based ones with identical language models. The active ingredient is judgment-free listening, not therapeutic framework.

Does therapist self-reference language predict weaker therapeutic alliance?

High frequency of therapist 'I' usage correlates with lower patient-reported alliance and reduced trusting behavior in validated behavioral tasks. Patient non-fluency markers like filler pauses, conversely, signal relaxed communication and stronger alliance.

Can conversational AI safely take patient histories without supervision?

A single-arm study found that AMIE, a conversational AI system, conducted real clinical histories from 100 patients without requiring a single safety intervention by human supervisors. Patient attitudes toward AI improved after the interaction, though management plans trailed physicians on practicality and cost.

Show all 6 sources
Can language models match therapist empathy in real conversations?

Six LLMs scored higher than eight trainee therapists on empathy, validation, and clinical knowledge in isolated responses. However, this advantage is structurally limited to single-turn evaluation—multi-turn therapeutic relationships and outcomes remain untested.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.