Curated summary
Exploring the feasibility of conversational diagnostic AI in a real-world clinical study
The study evaluated Google’s conversational medical AI, AMIE, in a real-world primary care workflow rather than simulated cases. In a prospective, IRB-approved study at Beth Israel Deaconess Medical Center, AMIE conducted supervised pre-visit history-taking with 100 patients. Results suggested that supervised deployment was feasible and conversationally safe, while AMIE’s diagnostic and management-plan quality was broadly comparable to that of primary care physicians, with physicians performing better on practicality and cost effectiveness.
Study Design and Clinical Workflow
- Patients with new, non-emergency, episodic complaints used AMIE through a secure web link before an in-person or telehealth appointment.
- A physician supervised each AI-patient interaction through live video and screen-sharing.
- AMIE produced a transcript and summary for the patient’s primary care physician.
- Independent clinical evaluators assessed:
- The quality of the AMIE conversation
- AMIE’s differential diagnoses
- AMIE’s management plans
- Comparable outputs from physicians
- The study was prospective, single-center, single-arm, pre-registered, and IRB approved.
Participants
- 100 adults completed the AMIE interaction.
- 98 attended their scheduled primary care appointments.
- Participants represented varied ages, racial and ethnic groups, health literacy, technology literacy, and prior chatbot experience.
- Compared with all 1,452 urgent care visits during the study period, participants tended to be younger, although the sample reflected the broader population’s female and white demographic skew.
Safety Oversight
- Human supervisors could stop an interaction if they observed:
- Immediate risk of harm to the patient or others
- Significant emotional distress related to the AI interaction
- Potential clinical harm
- A patient’s explicit request to end the session
- No safety stops were required across the study.
- The authors interpret this as evidence that supervised AMIE interactions were conversationally safe in this setting.
Clinical Reasoning Performance
- Three independent clinical evaluators reviewed each case using blinded, randomized assessments.
- AMIE and physicians showed similar overall quality for:
- Differential diagnoses
- Management plans
- Management-plan appropriateness and safety
- Physicians performed better on the practicality and cost effectiveness of management plans.
- AMIE’s differential-diagnosis accuracy was reported as high, including cases where the final diagnosis was confirmed through diagnostic testing.
Patient and Clinician Experience
- The study measured trust, perceptions, and acceptance among both patients and clinicians.
- Patient trust in AI increased after interacting with AMIE.
- Overall findings indicated that the system was well received within the supervised pre-visit workflow.
The study supports cautious, supervised testing of conversational diagnostic AI in clinical environments. It does not establish that AMIE can independently replace clinicians; rather, it suggests that pre-visit information gathering may be a practical early use case, provided rigorous oversight, safety protocols, and further evaluation in larger and more diverse settings.
Related reading
Continue with another curated summary.
SymptomAI: Towards a conversational AI agent for everyday symptom assessment
Read originalHow AI Has Changed the Product Design Process
Read originalAdvancing AMIE towards expert-level audio-visual clinical consultations
Read originalResearch into how AI can help users understand skin conditions
Read original