Google's AMIE study reached a new venue on October 8, 2026: The Lancet's main journal, a first for Google. AMIE is Google's conversational diagnostic research system, and this time it did not sit a simulated exam. At Beth Israel Deaconess Medical Center's primary care clinic, it took histories from real patients before their urgent care visits.
How the study ran
Ninety-eight patients chatted with AMIE ahead of their appointments while supervising physicians monitored every interaction in real time, with predefined safety criteria for stopping a conversation. Not one conversation had to be interrupted. AMIE's differential diagnosis lists, checked against chart review eight weeks later, included the patient's final diagnosis in 90% of cases; narrowed to the top three, the hit rate was still 75%. For the doctors, the output was more than decoration: clinicians reported that AMIE's summaries helped them prepare in 75% of cases and influenced their approach to care in more than half.
The weaknesses are on the record too
The paper does not present AMIE as a doctor replacement. In blinded assessment, its management plans lost to primary care physicians on practicality and cost-effectiveness: listing plausible diseases is the easier part; sequencing affordable, workable next steps for a specific patient is harder. The study's limits are plain as well — a single center, text-only interaction, roughly a hundred patients, and no control group. Google's own account stresses that larger clinical trials are still needed before patient-facing AI can be judged at scale.
Why this step stands out
Most medical AI report cards of recent years come from question banks and simulated consultations. However high those scores run, they cannot answer real-clinic questions: patients describe symptoms vaguely, contradict themselves, and a named physician answers for every recommendation. The value here is not the 90% figure itself, but that supervision arrangements, stop criteria, and genuine clinician feedback were published together in a top journal's prospective study. For hospital decision-makers, it adds one sharp question to any medical AI pitch: has this system been tested inside a real workflow, with a human able to pull the plug at any moment? Without that step, a beautiful benchmark score is still only a laboratory result.