Clinical Reasoning

2 posts

google3 min readCurated summary

Exploring the feasibility of conversational diagnostic AI in a real-world clinical study

The study evaluated Google’s conversational medical AI, AMIE, in a real-world primary care workflow rather than simulated cases. In a prospective, IRB-approved study at Beth Israel Deaconess Medical Center, AMIE conducted supervised pre-visit history-taking with 100 patients. Results suggested that supervised deployment was feasible and conversationally safe, while AMIE’s diagnostic and management-plan quality was broadly comparable to that of primary care physicians, with physicians performing better on practicality and cost effectiveness. ## Study Design and Clinical Workflow - Patients with new, non-emergency, episodic complaints used AMIE through a secure web link before an in-person or telehealth appointment. - A physician supervised each AI-patient interaction through live video and screen-sharing. - AMIE produced a transcript and summary for the patient’s primary care physician. - Independent clinical evaluators assessed: - The quality of the AMIE conversation - AMIE’s differential diagnoses - AMIE’s management plans - Comparable outputs from physicians - The study was prospective, single-center, single-arm, pre-registered, and IRB approved. ## Participants - 100 adults completed the AMIE interaction. - 98 attended their scheduled primary care appointments. - Participants represented varied ages, racial and ethnic groups, health literacy, technology literacy, and prior chatbot experience. - Compared with all 1,452 urgent care visits during the study period, participants tended to be younger, although the sample reflected the broader population’s female and white demographic skew. ## Safety Oversight - Human supervisors could stop an interaction if they observed: - Immediate risk of harm to the patient or others - Significant emotional distress related to the AI interaction - Potential clinical harm - A patient’s explicit request to end the session - No safety stops were required across the study. - The authors interpret this as evidence that supervised AMIE interactions were conversationally safe in this setting. ## Clinical Reasoning Performance - Three independent clinical evaluators reviewed each case using blinded, randomized assessments. - AMIE and physicians showed similar overall quality for: - Differential diagnoses - Management plans - Management-plan appropriateness and safety - Physicians performed better on the practicality and cost effectiveness of management plans. - AMIE’s differential-diagnosis accuracy was reported as high, including cases where the final diagnosis was confirmed through diagnostic testing. ## Patient and Clinician Experience - The study measured trust, perceptions, and acceptance among both patients and clinicians. - Patient trust in AI increased after interacting with AMIE. - Overall findings indicated that the system was well received within the supervised pre-visit workflow. The study supports cautious, supervised testing of conversational diagnostic AI in clinical environments. It does not establish that AMIE can independently replace clinicians; rather, it suggests that pre-visit information gathering may be a practical early use case, provided rigorous oversight, safety protocols, and further evaluation in larger and more diverse settings.

Read original(opens in new tab)
googleOriginal article

How Google’s AI can help transform health professions education (opens in new tab)

To address a projected global deficit of 11 million healthcare workers by 2030, Google Research is exploring how generative AI can provide personalized, competency-based education for medical professionals. By combining qualitative user-centered design with quantitative benchmarking of the pedagogically fine-tuned LearnLM model, researchers have demonstrated that AI can effectively mimic the behaviors of high-quality human tutors. The studies conclude that specialized models, now integrated into Gemini 2.5 Pro, can significantly enhance clinical reasoning and adapt to the individual learning styles of medical students. ## Learner-Centered Design and Participatory Research * Researchers conducted interdisciplinary co-design workshops featuring medical students, clinicians, and AI researchers to identify specific educational needs. * The team developed a rapid prototype of an AI tutor designed to guide learners through clinical reasoning exercises anchored in synthetic clinical vignettes. * Qualitative feedback from medical residents and students highlighted a demand for "preceptor-like" behaviors, such as the ability to manage cognitive load, provide constructive feedback, and encourage active reflection. * Analysis revealed that learners specifically value AI tools that can identify and bridge individual knowledge gaps rather than providing generic information. ## Quantitative Benchmarking via LearnLM * The study utilized LearnLM, a version of Gemini fine-tuned specifically for educational pedagogy, and compared its performance against Gemini 1.5 Pro. * Evaluations were conducted using 50 synthetic scenarios covering a spectrum of medical education, ranging from preclinical topics like platelet activation to clinical subjects such as neonatal jaundice. * Medical students engaged in 290 role-playing conversations, which were then evaluated based on four primary metrics: overall experience, meeting learning needs, enjoyability, and understandability. * Physician educators performed blinded reviews of conversation transcripts to assess whether the AI adhered to medical education standards and core competencies. ## Pedagogical Performance and Expert Evaluation * LearnLM was consistently rated higher than the base model by both students and educators, with experts noting it behaved "more like a very good human tutor." * The fine-tuned model demonstrated a superior ability to maintain a conversation plan and use grounding materials to provide accurate, context-aware instruction. * Findings suggest that pedagogical fine-tuning is essential for AI to move beyond simple fact-delivery and toward true interactive tutoring. * These specialized learning capabilities have been transitioned from the research phase into Gemini 2.5 Pro to support broader educational applications. By integrating these specialized AI behaviors into medical training pipelines, institutions can provide scalable, individualized support to students. The transition of LearnLM’s pedagogical features into Gemini 2.5 Pro provides a practical framework for developers to create tools that not only provide medical information but actively foster the critical thinking skills required for clinical practice.