INQUIRING LINE

Does the packaging of an AI tool matter as much as its accuracy when doctors decide whether to use it?

What role does interface design play in clinician adoption of AI tools?

This explores whether the way an AI tool is presented and structured, not just how accurate it is, shapes whether clinicians actually take it up. The corpus has no direct study of clinician adoption, so the answer pieces it together from related work on medium, structure and scaffolding.


This reads the question as asking whether the way an AI tool is presented and structured, rather than how accurate it is, decides whether clinicians use it. The collection has no study that tracks clinicians adopting or rejecting a tool because of its interface, so that gap should be clear up front. What it does have is a consistent pattern from nearby areas: the packaging often decides the outcome more than the model inside it does.

The strongest evidence comes from therapy. In a 15-day study, a robot and a structured worksheet reduced students' psychological distress, while a chatbot using the same language model did not Why do robots outperform chatbots in therapy despite identical language models?. Medical diagnosis shows the same thing from another angle. Wrapping o3 in an orchestration layer, a structured process that manages how the model gathers and weighs evidence, slightly raised its accuracy on hard NEJM cases and cut the cost per case by about 70%. The gains carried over to other model families, which points to the structure as the source rather than the model weights Can orchestration strategies boost diagnostic AI without better models?. If a clinical tool underperforms, the interface or workflow is a real suspect, not only the model.

Why might a chat box be the wrong interface for clinicians? One argument is that conversational design triggers people's lifelong communication habits, but the AI doesn't actually communicate the way a colleague does. The failures that follow feel like user error even though they come from the design Why do users fail with AI interfaces designed like conversations?. A related problem is that people often can't say exactly what they want up front. Their intent takes shape through back-and-forth, and models that only respond, rather than ask, miss that. Showing users model-generated options to choose from turns open-ended asking into picking from a list Why can't users articulate what they want from AI?. That suits busy clinicians, who are better at judging than at writing prompts. The corpus also argues that an AI's context is constantly shifting and partly hidden, so users can't learn it the way they learn a fixed software screen. On this view, designing what the model sees matters as much as designing what the user sees How does AI context differ from conventional software context?.

Here is the surprising part. Output quality may not be what holds adoption back. In blind ratings, clinicians found GPT-4's written advice as scientifically sound as expert advice and more emotionally empathetic, and they could tell the two apart only at chance level Can clinicians tell GPT-4 advice apart from expert advice?. If the text is indistinguishable, the barriers lie elsewhere: who is accountable, and how the tool fits into the work. On the patient side, the barriers people report are about uniqueness, perceived competence and accountability, and they exist regardless of what the AI can actually do Why do patients distrust medical AI systems?. People also judge AI partners mainly on perceived competence, which accounts for about half of their overall impression How do users mentally model dialogue agent partners?. Perceived competence is something an interface can build up or wear down.

The most concrete picture of a clinician-facing design is the 'AI supervisor' model. A system transcribes a therapy session, scores the therapist-patient working alliance turn by turn, and suggests the next topic in real time [[rl-based-topic-recommendation-systems-can-serve-as-real-time-ai-supervisors-for][working-alliance-can-be-computationally-inferred-from-session-transcripts-at-tur]]. This is a very different interface from a chatbot. It sits alongside the clinician's own work and offers suggestions without taking over the conversation. Whether clinicians welcome this kind of tool, or find it intrusive, is the open question the corpus doesn't yet answer.


Sources 10 notes

Why do robots outperform chatbots in therapy despite identical language models?

A 15-day study with 38 students found that robots and worksheets significantly reduced psychological distress while a chatbot using the same LLM did not. The active ingredient was the medium—social presence and structured format—not language capability.

Can orchestration strategies boost diagnostic AI without better models?

On 304 NEJM cases, MAI-DxO orchestration achieved 79.9% accuracy at $2,397 per case versus 78.6% at $7,850 for o3 alone. The gains transferred across model families, suggesting the benefit comes from the scaffold, not model weights.

Why do users fail with AI interfaces designed like conversations?

AI interfaces that use conversational design conventions trigger users' lifelong communication skills, but AI doesn't actually communicate. This mismatch causes interaction failures that feel like user error but originate in design.

Why can't users articulate what they want from AI?

Intent develops through interaction, not in isolation. Since AI models respond rather than probe, they miss opportunities to help users discover unarticulated requirements. Structured dialogue that presents model-generated options shifts the cognitive burden from open-ended envisioning to constrained evaluation.

How does AI context differ from conventional software context?

AI interactions operate on a substrate of constantly shifting context—prompt, history, retrieved data, hidden state—that users cannot internalize like traditional UIs. This structural mutability demands a new design discipline centered on context engineering rather than interface design.

Show all 10 sources
Can clinicians tell GPT-4 advice apart from expert advice?

Blinded clinician ratings of 104 response pairs found GPT-4 advice favored on emotional empathy, with no significant differences in scientific quality or cognitive empathy. Clinicians identified the source at chance level (45% accuracy), suggesting the two were indistinguishable in written form.

Why do patients distrust medical AI systems?

Research identifies three distinct user-side barriers: patients perceive AI as unable to address their unique needs, believe it performs worse than human providers, and see it as harder to hold accountable. These barriers exist independent of actual AI capability.

How do users mentally model dialogue agent partners?

The Partner Modelling Questionnaire reveals that perceived competence dominates user impressions (49% of variance), followed by human-likeness (32%) and communicative flexibility (19%). This three-factor structure reflects how people evaluate dialogue partners against both functional and social standards.

Can reinforcement learning optimize therapy dialogue in real time?

R2D2 demonstrates that RL agents trained on multi-objective working alliance scores can generate disorder-specific policies that recommend treatment strategies in real time. The system operates as an AI supervisor, transcribing sessions and recommending next topics based on task, bond, and goal alignment.

Can we measure therapist-patient alliance from dialogue turns in real time?

COMPASS maps dialogue turns onto WAI embeddings to produce 36-dimensional alliance scores per turn. Anxiety and depression show convergence in alliance metrics over time, while suicidality shows persistent misalignment between patient and therapist.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.