The future of federally funded research at Harvard Medical School — supported by taxpayers and done in service to humanity — remains uncertain. Learn more.
Scientific advancements — including understanding health conditions and developing new treatments — hinge on data and its implications. Statistician Alex Luedtke is developing new methods to make data analysis as efficient and effective as possible with the goal of improving personalized treatment recommendations in mental health and other fields.
Luedtke joined the faculty in the Blavatnik Institute at Harvard Medical School last summer as professor of health care policy. His work combines statistical methodology with machine-learning techniques to draw cause-and-effect conclusions from messy biomedical data and speed discovery.
“I really enjoyed the mathematical sciences, but I wanted to do something with them that felt meaningful,” said Luedtke, who previously held faculty appointments at the University of Washington and Fred Hutch.
Luedtke has received an Emerging Leader Award from the Committee of Presidents of Statistical Societies, the Mortimer Spiegelman Award, and a National Institutes of Health Director’s New Innovator Award.
Harvard Medicine News spoke with Luedtke about his collaborations at Harvard and his ongoing efforts to build new approaches to data analysis that could facilitate personalized medicine and improve patient outcomes.
Harvard Medicine News: Tell us a little bit about the focus of your research.
Alex Luedtke: I’m a statistician. Causal inference and machine learning are my two main areas, so broadly speaking, I try to pull cause-and-effect answers out of real-world data.
A lot of my work focuses on what we call efficiency theory, which is a fancy name for using data as well as possible. We’re trying to get the most precise answers possible from a limited amount of data.
One specific question I’m working on, which I started back during my doctoral dissertation at the University of California, Berkeley, is the extent to which individualizing treatment decisions can help improve outcomes for people. Based on patient characteristics, how can we best decide whether we should give them treatment A, B, or C? In my previous role, I did this in infectious disease settings, but increasingly nowadays I’m doing it in mental health with collaborators here at Harvard.
HMNews: What infectious diseases were you working on?
Luedtke: I worked on vaccines trials, both for HIV and then also COVID. I served as a study statistician, which meant that I would monitor the trials as they went and also determine what statistical methods would be used to analyze the data from those trials.
Clinical trials are a classic use case for efficient statistical methods. It’s very expensive to recruit new people into a trial and to run a trial longer to keep accruing more information about participants. And ethically, you want to be able to terminate a trial if a vaccine isn’t working. So it’s best for everyone involved to figure out whether a vaccine is effective as quickly as possible.
HMNews: What do your current mental health collaborations look like?
Luedtke: I actually started collaborating with [McNeil Family Professor of Health Care Policy] Ron Kessler, who is in my current department, back in 2016. In our collaborations, we work a lot on figuring out whether individualization will help for treating mental health disorders. We’ve looked at depression and at schizophrenia. We take the statistical methods that I develop for treatment individualization and apply them directly to his datasets, most of which are from observational, non-randomized trials.
That’s been a lot of fun. When statisticians develop new statistical methods, we always develop them in the cleanest possible setting where we get to see the outcomes on everyone, and we have a random sample from the population we care about. The moment we start touching real data, we don’t get that. Patients drop out of the study or are just lost in the registry. And sometimes the patients who drop out are different than other patients in some meaningful way that actually has something to do with whether the treatment works.
And so a lot of my work with Ron is looking at each particular dataset and asking: How was the data generated? Why are data missing? What do we have to do statistically to make sure that we can develop unbiased conclusions at the end? Ultimately, that analysis helps us determine if it’s worth individualizing treatments for these mental health disorders and make recommendations to help physicians provide the best care for their patients.
HMNews: So how do machine learning and AI come into it?
Luedtke: AI looks at prediction problems — trying to predict what will happen in the world we have now. Causal inference asks: If we were to go in and change something in the world — introduce a new treatment or make a treatment more widely available — what will happen?
The challenge is that we’re changing something. Even if we have data from the world as it is now, we won’t have data from the world as we want it to be. One of the projects I’ve looked at is training these generative AI tools to essentially create new data for a world in which we change the treatment.
It’s similar to how we train existing generative AI models, but with the extra challenge that we want to determine cause and effect and we don’t have existing data from that scenario.
HMNews: I know you’re just getting started here, but over the coming years — or decades — what are you hoping to accomplish? Where do you want to move the needle forward in your field?
Luedtke: Historically, statisticians derive new methods by sitting down with a pen and paper and writing out how we think the data were generated and then spending a bunch of time building an estimator — a new way to analyze that data. Then methodological statisticians work to prove those results, essentially certifying the analysis strategy. Once that’s done, we can go out and use it on real data. This is a lot of what statistics has been for a hundred years.
I’m very interested in speeding up this process. If you tell me that I’m going to get a new dataset that was generated in a certain way with specific quirks, I’d rather avoid sitting down with the pen and paper at all.
This isn’t a new interest for me. A few years ago, we published a paper in Science Advances on getting a machine to work out a statistical procedure on its own, rather than a person deriving it by hand. Back then we had to train a model from scratch for every new statistical setting we wanted to handle.
Now AI models can carry what they learn in one setting into settings they’ve never seen. This ability is behind the scientific foundation models people are building in other fields, where a single model generates predictions across a whole scientific domain. I’m trying to build something similar for statistics: give the model a new statistical setting, and it works out the estimator. We have to be very careful about how we build it, but it could let us skip the slowest and most error-prone parts of the current process.
If we can do that, we can speed up how quickly we are able to draw scientific conclusions from data and translate them into treatments that can start helping patients.
This interview was edited for length and clarity.