Earlier this summer, an international team led by biomedical informatics researchers at Harvard Medical School released the first public dataset that connects descriptive radiological information, such as “3mm nodule in the lower left lobe,” to precise locations in chest CT scans.
AI models trained on this dataset could analyze CT scans and communicate results more effectively with clinicians, reporting findings in descriptive language and linking those descriptions to specific pixels within three-dimensional scans.
The work, published June 18 in the New England Journal of Medicine AI, addresses a critical gap in medical AI by making radiology findings clearer and easier to confirm, the researchers said.
“The finding becomes something a clinician can verify instead of something they have to take on faith from an AI tool,” said senior author Pranav Rajpurkar, associate professor of biomedical informatics in the Blavatnik Institute at HMS.
The researchers made the dataset, ReXGroundingCT, accessible to the larger scientific community leading up to an official challenge at the International Conference on Medical Image Computing and Computer Assisted Intervention, which concluded on Sept. 27. The challenge drew 53 teams from 15 countries testing different approaches to building models that can link natural language to locations in a CT scan, a process known as grounding.
“This turns grounding from a nice idea into a benchmark people can build against,” Rajpurkar said. “The challenge produced state-of-the-art solutions for this task, pushing results well past where the field stood when we started.”
Grounding radiology findings
AI models have the potential to help radiologists analyze medical scans quickly and thoroughly, but clinicians need to be able to confirm their results. This is particularly challenging with CT scans, which generate complex 3D images.
Radiologists visualize these images as a series of 2D slices and create reports that describe the locations of anomalies. For an AI model to do the same, it needs to pair clear and accurate descriptions with visual markers that show clinicians exactly where to look across hundreds of 2D slices.
“Datasets like ours will enable training models that can not only output a report but also show you exactly where it thinks these findings exist, so it becomes much easier for the clinician to verify each finding,” said Mohammed Baharoon, a Harvard Kenneth C. Griffin Graduate School of Arts and Sciences PhD student in Rajpurkar’s lab and first author on the paper.
Over the course of a year, Rajpurkar, Baharoon, and their colleagues coordinated with medical annotators, medical students, and radiologists to manually annotate more than 16,000 anomalies across 3,142 scans. Each annotation connected language in a report to specific locations within a CT scan and was double-checked by a board-certified radiologist.
“My biggest hope is that this trains a generation of models that don’t just describe a finding but actually show where it is,” Rajpurkar said. “A model shouldn’t only say, ‘there’s a nodule.’ It should put a marker on the exact voxel it’s talking about, so every statement comes with its evidence.”
Applications for medical students and patients
The researchers see additional uses for the dataset in medical education. Rajpurkar and Baharoon previously designed an AI-powered platform for radiology training based on a dataset of annotated X-rays. The platform helps trainees learn to write reports and localize findings.
With the ReXGroundingCT dataset, the researchers hope to create a training module for CT scans, which are significantly more complicated than X-rays.
The dataset could be used for tools that improve patient experience as well, Rajpurkar said. His lab has already built an AI system that generates video explanations of radiology findings to help patients understand their diagnostic results.
“Grounding enables tools that show patients exactly where on their own scan a finding sits and translate the surrounding jargon into plain language,” he said.
Authorship, funding, disclosures
Luyang Luo is co-first author of this work. Additional authors include Michael Moritz, Abhinav Kumar, Sung Eun Kim, Xiaoman Zhang, Miao Zhu, Mahmoud H. Alabbad, Maha S. Alhazmi, Neel P. Mistry, Lucas Bijnens, Kent R. Kleinschmidt, Brady Chrisler, Sathvik Suryadevara, Sri Sai Dinesh Jaliparthi, Noah M. Prudlo, Mark D. Marino, Jeremy Palacio, Rithvik Akula, Di Zhou, Hong-Yu Zhou, Ibrahim E. Hamamci, Scott J. Adams, and Hassan R. Alomaish.
This research was supported by the Harvard Medical School, Seoul National University Hospital, and Seoul National University College of Medicine collaborative research program, sponsored by the latter (project no. 8621382-01); the NVIDIA Academic Grant Program; and the HMS Dean’s Innovation Award for the Use of Artificial Intelligence in Education, Research, and Administration. Bijnens was supported by the Junior Orsi and Barco Fellowships.
Moritz is the chief medical information officer for a2z Radiology AI. Rajpurkar is a co-founder of a2z Radiology AI.