Can AI safety nets reduce errors in live clinical practice?
A study of 39,849 visits at Nairobi primary care clinics tested whether LLM decision support tools could help clinicians make fewer diagnostic and treatment errors in routine care, and what conditions made the tool effective.
This quality-improvement study, run with Penda Health's primary care clinics in Nairobi, compares 39,849 visits by clinicians with and without access to AI Consult, an LLM tool that serves as a safety net for clinicians. Independent physicians reviewing documentation rated the AI group's visits as having 16% fewer diagnostic errors and 13% fewer treatment errors. A clinician survey reports that every AI-group respondent said the tool improved the care they could deliver, and 75% called the effect "substantial." The authors project that the tool would avert diagnostic errors in 22,000 visits and treatment errors in 29,000 visits a year at Penda alone, but that is their projection, not a measurement. Their summary is that "these results required a clinical workflow-aligned AI Consult implementation and active deployment to encourage clinician uptake."
The paper names two conditions beyond model capability. The first is implementation. The earlier AI Consult v1 required clinicians to stop mid-visit and request feedback, reached about 60% of visits, and left many correct suggestions unheeded. The studied version fires asynchronously at EMR decision points behind a green, yellow or red interface, "explicitly designed to minimize cognitive burden and preserve clinician autonomy." The second is active deployment: the effect was significantly larger in the main period, when deployment strategies ran, than in the induction period. The authors also read a learning effect into a 10–15 point drop in visits that "started red" for treatments in the AI group, while the non-AI group stayed flat. Their conclusion lists "capable models, which are now widely available" as only one of three components.
Against the nearest notes, the paper is a live-care counterpart to offline evaluation. Can language models match therapist empathy in real conversations? rests on isolated responses. This study measures errors in routine visits, which the authors say offline work cannot capture: "real-world patient diversity, designing for and learning from clinician workflows, and deployment towards successful clinician uptake." The workflow emphasis echoes Can clinical experts teach LLMs to annotate complex medical concepts?, which also places clinical-LLM friction in the interface between expert and tool. The direction is reversed, though: that excerpt reports barriers, while this one reports what made a deployed tool work. On the uptake side, Why do patients distrust medical AI systems? concerns patient attitudes. The uptake problem here is clinicians not acting on correct advice.
The excerpt does not establish that the workflow and deployment choices caused the error reduction rather than the model alone. The comparison was access versus no access under a cluster-assigned design, and the studied implementation was never compared with a non-aligned version on error rates. v1 was assessed only for adoption and a 100-case internal safety audit, so "required" is the authors' inference from design rather than a tested contrast. Patient-reported outcomes showed no significant difference, which the authors attribute possibly to measurement sensitivity, a 40% response rate, and a short follow-up. The error ratings are physician judgments of documentation, not observed patient harm, and several authors are affiliated with Penda, as the introduction states. The supportable reading is narrow: in one large Kenyan network, a documentation-based error reduction appeared alongside these implementation and deployment choices. It does not yet show better patient outcomes or transfer to other health systems. The authors call their work "a first step."
Inquiring lines that read this note 6
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What prevents LLMs from applying their reasoning knowledge to improve outputs? How do clinicians calibrate trust in AI medical recommendations?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can language models match therapist empathy in real conversations?
Do LLMs' high empathy scores on isolated responses translate to therapeutic skill in actual ongoing treatment? This explores whether single-turn advantage predicts real-world therapeutic performance.
offline comparison on isolated responses; this study instead measures errors in live routine visits, which the authors say offline work cannot capture
-
Can clinical experts teach LLMs to annotate complex medical concepts?
Clinical experts can manually identify complex medical concepts in patient notes, but transferring that expertise to LLM-based extraction systems proves difficult. Understanding where this transfer breaks down could improve how AI tools support expert workflows.
parallel workflow-side account of clinical-LLM friction, but reports barriers where this study reports what made a deployed tool work
-
Why do patients distrust medical AI systems?
Explores the psychological barriers that make patients reluctant to adopt medical AI, beyond whether the technology actually works. Understanding these barriers is critical for designing AI systems patients will actually use.
patient-side adoption barriers; here the uptake problem is clinicians not acting on correct advice in v1
-
Why do LLMs fail when users interact with them?
Standard benchmarks show LLMs excel at medical diagnosis alone, yet real users get no benefit. This explores where the breakdown happens between model capability and human decision-making.
qualifies: LLM assistance left lay public users no better than controls in a 1,298-person UK trial, unlike these clinician gains
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- AI-based Clinical Decision Support for Primary Care: A Real-World Study
- Clinical knowledge in LLMs does not translate to human interactions
- A prospective clinical feasibility study of a conversational diagnostic AI in an ambulatory primary care clinic
- Capabilities of Gemini Models in Medicine
- Towards Accurate Differential Diagnosis with Large Language Models
- Superhuman performance of a large language model on the reasoning tasks of a physician
- Sequential Diagnosis with Language Models
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
Original note title
LLM decision support cut errors in live Nairobi primary care — the study says the gains required workflow-aligned implementation and active deployment