INQUIRING LINE

Training a diagnostic AI on AI-played patients looks promising, but the evidence hasn't pinned down how much that practice helps.

How much does self-play training with LLM-simulated patients actually improve diagnostic accuracy?

This explores whether training a diagnostic AI by having it practice on AI-played patients (self-play) actually makes it a better diagnostician, and by how much. The corpus can't give a clean number, but it explains why that number is hard to pin down.


This explores whether letting a diagnostic AI practice against LLM-simulated patients measurably improves its diagnoses. The short answer: the collection has no study that isolates self-play's contribution, with and without it. What it does have shows where the gains appear, where they disappear, and why practicing against simulated patients is a weaker training signal than it looks.

The headline results are strong. AMIE, an LLM-based diagnostic system, beat primary care physicians on 28 of 32 specialist-rated measures across 149 simulated consultations Can an AI system diagnose better than primary care doctors?. A separate study found an LLM outperforming hundreds of physicians on differential diagnosis, triage and management Can language models reason better than physicians at diagnosis?. Two caveats matter. AMIE's lead came mainly from reasoning over the information it had gathered, not from asking patients better questions. Yet asking questions is the skill that practice conversations with simulated patients should build most directly. AMIE was also evaluated in simulated, text-only consultations, so the test looks a lot like the training setup.

The weak point is the simulated patient itself. LLM user simulators drift away from their own goals over a multi-turn conversation. That drift corrupts the training signal, and one line of work had to break each simulated user's goals into separately tracked parts to keep it under control Why do LLM user simulators fail to track their own goals?. Realism has to be built deliberately. In therapy training, patients built on structured cognitive models were rated more realistic than plain GPT-4 role-play Can structured cognitive models improve LLM patient simulations for therapy training?. In other domains, simulators need explicit profile and intent variables to pass as realistic Can controlled latent variables make LLM user simulators realistic?. A model that trains on its own generated data also runs into error avalanching: small mistakes grow over two or three rounds, and improvement stalls at a level set by how well outputs are checked, not by what the model can actually do How quickly do errors compound during model self-training?.

The part you might not expect: even a perfect diagnostic model may not help real patients much. In a trial of 1,298 people, LLMs working alone named the right condition 94.9% of the time. People using those same LLMs got it right only 34.5% of the time, no better than people without them Why do LLMs fail when users interact with them?. Simulated patients behave like idealized partners and real people do not, so self-play may sharpen exactly the skill that matters least in practice. Clinicians are a different case. Doctors given LLM access reached 51.7% top-10 differential accuracy versus 36.1% with search alone, mostly because the model widened the list of possibilities they considered Does LLM assistance help clinicians build better differentials?. Confidence is another concern. In specialized clinical tasks, models stay overconfident even when their accuracy is low Why do language models fail confidently in specialized domains?, and practice against agreeable simulated patients is unlikely to correct that.

Overall, self-play has produced diagnostic systems that score very well in simulated consultations. The collection can't tell you how much of that comes from the self-play itself. It does give a reason to doubt that simulation scores carry over to real patients: the measured gains appear in settings that resemble the simulator, and they shrink once real people are part of the conversation.


Sources 9 notes

Can an AI system diagnose better than primary care doctors?

An LLM-based diagnostic system called AMIE exceeded primary care physician performance in text-based simulated consultations across 149 case scenarios, scoring higher on 28 of 32 specialist-rated dimensions. The advantage lay in inference from gathered information rather than in eliciting history.

Can language models reason better than physicians at diagnosis?

In physician-adjudicated vignette experiments and an emergency room study, a large language model outperformed hundreds of physicians on differential diagnosis, reasoning, triage, and clinical management tasks across multiple touchpoints.

Why do LLM user simulators fail to track their own goals?

The UGST framework breaks user goals into profile, policy, task, requirements, and preferences—each with explicit status tracking. A three-stage method (steering, SFT, GRPO) progressively internalizes goal alignment, reducing the misalignment that corrupts RL training signals.

Can structured cognitive models improve LLM patient simulations for therapy training?

PATIENT-Ψ integrates 106 Beck CCD-based cognitive models with LLMs to simulate patients with specific maladaptive patterns. Expert evaluators rated the fidelity higher than GPT-4, particularly for maladaptive cognitions and conversational authenticity.

Can controlled latent variables make LLM user simulators realistic?

RecLLM demonstrates that conditioning an LLM simulator on session-level (user profile) and turn-level (user intent) latent variables produces synthetic conversations measurable as realistic via crowdsource discrimination, discriminator models, and classifier-ensemble distribution matching.

Show all 9 sources
How quickly do errors compound during model self-training?

Small inaccuracies in model-generated training data amplify rapidly across iterations, degrading performance unless self-consistency checks filter outputs. The effect stalls improvement within a few steps, setting an error floor based on verification quality rather than actual capability.

Why do LLMs fail when users interact with them?

A trial of 1,298 UK participants found GPT-4o, Llama 3, and Command R+ scored 94.9% accuracy alone but users achieved only 34.5% when identifying conditions, no better than controls. The gap lies in how users interpret and act on model suggestions, not model knowledge.

Does LLM assistance help clinicians build better differentials?

In a study of 20 clinicians on 302 NEJM cases, those with LLM access achieved 51.7% top-10 accuracy versus 36.1% without it. The authors attribute the gain to the LLM's wider differential scope, making lists more comprehensive.

Why do language models fail confidently in specialized domains?

LLMs trained on general text lack sufficient exposure to domain-specific examples, leading to low accuracy paired with high confidence in clinical NLI tasks. Prompting techniques that improved general performance fail to reduce overconfidence in specialized domains.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.