The future of federally funded research at Harvard Medical School — supported by taxpayers and done in service to humanity — remains uncertain. Learn more.
Researchers have just completed one of the largest-yet studies comparing artificial intelligence and physicians across a wide range of clinical reasoning tasks, evaluating whether an AI system could do what physicians do every day: review a messy patient chart and decide what to do next.
A large language model (LLM) outperformed physicians across many of these tasks, including making emergency-room decisions based on the available information, identifying likely diagnoses, and choosing the next steps in management, a team led by physicians and computer scientists at Harvard Medical School and Beth Israel Deaconess Medical Center reported April 30 in Science.
“We tested the AI model against virtually every benchmark, and it eclipsed both prior models and our physician baselines,” said co-senior author Arjun (Raj) Manrai, assistant professor of biomedical informatics in the Blavatnik institute at HMS and founding deputy editor of NEJM AI.
The results make the case that medical AI is ready to be studied the same way as all new medical interventions: through carefully controlled, rigorous, prospective clinical trials in real care settings.
Manrai noted that these trials are necessary to evaluate whether, how, and where such tools should be deployed in clinical care as aids to human practitioners.
The model’s performance also suggests that longstanding ways of testing medical AI may no longer capture the abilities of current systems — pointing to a possible turning point for the field.
“Models are increasingly capable. We used to evaluate models with multiple-choice tests; now they are consistently scoring close to 100 percent, and we can’t track progress anymore because we’re already at the ceiling,” said co-first author Peter Brodeur, HMS clinical fellow in medicine at Beth Israel Deaconess.
Testing medical AI in the real world
Incorporating standards first created in the 1950s to train and evaluate doctors, the researchers compared how an AI system performed against hundreds of clinicians. The comparisons included case study diagnostic challenges, reasoning exercises, and real emergency department cases.
In one experiment, the team tasked the LLM with evaluating patients at various points in a standard emergency department setting, ranging from early triage to later admission decisions. At each stage, the model was given only the information available at that point — drawn directly from actual electronic health records — and asked to generate likely diagnoses and recommend what should happen next.
“To better understand real-world performance, we needed to test performance early in the patient course, when clinical data is sparse,” said co-first author Thomas Buckley, a Harvard Kenneth C. Griffin Graduate School of Arts and Sciences doctoral student and Dunleavy Fellow in HMS’ AI in Medicine PhD track and member of the Manrai Lab.
Unlike in prior studies, the team did not smooth out the messiness of real‑world care before testing the model; the emergency department cases were presented exactly as they appeared in the electronic health record.
“We didn’t pre‑process the data at all,” said co-senior author Adam Rodman, HMS assistant professor of medicine at Beth Israel Deaconess, director of AI programs for the Carl J. Shapiro Center for Education and Research, and associate editor of NEJM AI.
At the early decision points in the real-world emergency department cases, the model matched or exceeded attending physicians in diagnostic accuracy — a result that surprised even the researchers.
“I thought it was going to be a fun experiment but that it wouldn’t work that well. That was not at all what happened,” Rodman said.
The researchers emphasized that their results do not suggest that AI systems are ready to practice medicine autonomously or that physicians can be removed from the diagnostic process.
“A model might get the top diagnosis right but also suggest unnecessary testing that could expose a patient to harm,” Brodeur said. “Humans should be the ultimate baseline when it comes to evaluating performance and safety.”
Authorship, funding, disclosures
Additional authors on the study are Zahir Kanjee, Ethan Goh, Evelyn Bin Ling, Priyank Jain, Stephanie Cabral, Raja-Elie Abdulnour, Adrian D. Haimovich, Jason A. Freed, Andrew Olson, Daniel J. Morgan, Jason Hom, Robert Gallo, Liam G. McCoy, Haadi Mombini, Christopher Lucas, Misha Fotoohi, Matthew Gwiazdon, Daniele Restifo, Daniel Restrepo, Eric Horvitz, and Jonathan Chen.
The research was supported by the National Institutes of Health (grants R01ES032470, 1R01AI17812101, UM1TR004921, U01NS134358), the Harvard Medical School Dean’s Innovation Award for Artificial Intelligence, the Macy Foundation (awards B25-15 and P25-04), the Stanford Bio-X Interdisciplinary Initiatives Seed Grants Program, and a Stanford RAISE Health Seed Grant 2024.
Rodman is a visiting researcher at Google DeepMind. Goh is employed by Microsoft. Chen is co-founder of Reaction Explorer LLC; a paid medical expert witness from Elite Experts; and received one-time honoraria or travel expenses for invited presentations by Insitro, General Reinsurance Corporation, AASCIF, and other industry conferences, academic institutions, and health systems. Kanjee discloses royalties from Oakstone Publishing and Wolters Kluwer. Olson discloses employment of his spouse by Exact Sciences. Abdulnour is employed by the Massachusetts Medical Society and has consulted for Lumeris.