Menu
An abstract illustration shows how three colorful, glowing balls generate a complex array of data points.
Image: Ihor Melnyk/iStock/Getty Images Plus

New AI Tool Predicts Risk of More Than 300 Diseases With Existing Patient Data

Machine learning reveals complex patterns that can guide care and research

Research 4 min read
By BETH DOUGHERTY | DANA-FARBER COMMUNICATIONS

At a glance

  • Researchers have developed the first known machine-learning-powered tool that can predict patient risk of hundreds of diseases based solely on patient health records and genetic profiles.

  • The model assesses complex relationships across multiple biological systems and medical specialties.

  • The tool could guide research and, if validated in the clinic, aid in fending off preventable illnesses, such as certain cancers and heart disease.

Harvard Medical School researchers at Dana-Farber Cancer Institute and Massachusetts General Hospital have created a machine-learning algorithm that can predict the likelihood of 348 distinct diseases for a given patient.

The algorithm, described in Nature July 15, is the first to make multi-disease predictions based only on routinely collected electronic health record data and knowledge about the patient’s genetic risks of disease.

The algorithm dynamically updates its risk assessments over time, and its predictive power improves as the patient ages.

Get more HMS news

“People are thinking about what their health is going to look like over the next few years, especially with increasing intervention options,” said co-senior author Alexander Gusev, HMS associate professor of medicine at Dana-Farber. “This tool offers a path toward improving the prediction of future diseases so doctors and patients can take action to try to prevent them.”

At this time, the tool has only been tested on past patient data. Future work is needed to assess it in clinical settings, as a research tool, and as a way to improve the design of clinical trials.

The team — led by Gusev; co-senior author Giovanni Parmigiani, Dana-Farber researcher and professor of biostatistics at the Harvard T.H. Chan School of Public Health; co-senior author Pradeep Natarajan, HMS professor of medicine at Mass General; and first author Sarah Urbut, HMS instructor in medicine at Mass General — combined models of probability with machine learning.

Artificial intelligence built on a foundation of human expertise

The researchers mathematically defined 20 biological “signatures,” or sets of biological trends, that have a high probability of initiating certain diseases. For example, high cholesterol in a patient’s history increases the probability of cardiovascular diseases. The signatures also account for genetic factors known to increase the likelihood of a disease.

These signatures are complex and overlapping. The risk of colon cancer, for instance, is associated with multiple signatures.

“We have painstakingly curated these signatures, which is a big differentiator [from other models]. In contrast to deep-learning approaches, which are typically ‘black boxes,’ our curated signatures capture the underlying biology in an interpretable way,” said Parmigiani. “These signatures are then the drivers of the model’s ability to make predictions.”

The team then trained an artificial intelligence-based model to predict the risk of a wide range of diseases based on the patient’s health history, viewed through the lens of their curated disease signatures and a learned understanding of the co-occurrence of diseases.

“There is a lot of useful information in a medical record, both over time and across different disease areas,” said Parmigiani. “That data would be difficult for a human to process in their head but tractable for a machine-learning model.”

Earlier, better predictions can improve preventive medicine

The team called their model Aladynoulli — a portmanteau of Aladdin, who had access to wish-granting genies, and mathematician Jacob Bernoulli, who defined the idea of dynamic predictions based on probabilities.

The model was trained and validated using three large biobanks including a total of over 683,000 patient records. The model provided more accurate 10-year predictions than three existing cardiovascular risk models — PCE, QRISK3, and PREVENT. It also outperformed the GAIL breast cancer risk model for one-year breast cancer predictions.

The model can also predict which patients will develop colorectal cancer in the coming year with high accuracy. The researchers suggested that flagging a high risk of imminent colorectal cancer could potentially help a primary care physician refer a patient for a colonoscopy, even if that patient is not yet eligible for screening based on current age-based guidelines.

By incorporating each patient’s underlying signature profiles, Aladynoulli could become an important tool to help power personalized medicine approaches. The tool could also improve patient care by breaking information out of silos based on medical specialty, said Parmigiani.

“This model is an innovative and potentially disruptive tool in the clinic because it could encourage more clinicians to think in a cross-disciplinary way,” he said.

A tool to improve patient care and to deepen research

Next steps for the team include using the model to learn more about why and how melanoma, a form of skin cancer, spreads to other organs. The team has partnered with medical oncologist David Liu, HMS assistant professor of medicine at Dana-Farber, to train Aladynoulli to identify different patterns of metastases, such as spreading early or late or widely or narrowly. The work could help identify previously unrecognized subtypes of the disease.

“Aladynoulli not only can help predict different trajectories of metastases but can also give us a biological understanding of why patients are progressing along these different trajectories,” said Gusev. “From a better understanding of the biology, we have the potential to identify more beneficial therapeutics.”

In addition, the team is working to expand the signatures to increase the accuracy and biological grounding of the risk predictions.

Adapted from a Dana-Farber news release.

Authorship, funding, disclosures

Additional authors include Yi Ding, Tetsushi Nakao, Satoshi Koyama, Anika Misra, Xilin Jiang, Achyutha Harish, Leslie Gaffney, Whitney E. Hornsby, and Jordan W. Smoller.

This work was supported by the National Institutes of Health (grants R01HL127564, R00HL165024, 1K08HL183784, U01HG011719), American Heart Association (Career Development Award 25CDA1444806), Burroughs Wellcome Fund (award 1360373), Wellcome Early Career Award (227566/Z/23/Z), and a gift from the Dobson family.