About Me: I’m an Assistant Professor of Computer Science at Johns Hopkins, and a part-time Member of Technical Staff at Abridge. Previously, I was a postdoc at Carnegie Mellon University with Zack Lipton, and obtained my PhD in Computer Science at MIT with David Sontag.

News

Research Overview

My group develops methods for principled and efficient evaluation and monitoring of AI systems, with an emphasis on applications in healthcare. We draw on a broad methodological toolkit spanning statistics, causal inference, and machine learning, and collaborate closely with clinical researchers in medicine, nursing, and public health. Our work can be seen as covering three areas, detailed below. A few representative papers are listed under each; see the Papers page for the full list.

Methods for scalable, automated assessment of AI systems

It is often challenging to scalably define what “good” looks like for a generative AI system in a high-expertise, non-verifiable domain like healthcare. Our group develops new methods and interrogates existing methods for trying to do so at scale, leveraging e.g., existing clinician documentation (in the case of e.g., radiology report generation), or expert-written rubrics.

Evaluating Rubric Generation with Interventional Transfer
Erik Skalnes, Layne C. Price, Raviteja Anantha, Michael Oberst
preprintcode

Reporting Practice Matters: The Impact of Reference Choice on Chest X-ray Report Evaluation
Daniel P. Jeong, Charles Q. Li, Hossein Hosseiny, Nitya M. Bhalla, Fatma Uyar Morency, Pradeep Ravikumar, Zachary C. Lipton, Michael Oberst
Neural Information Processing Systems, Evaluation and Datasets Track (NeurIPS), 2026
paper

Methods for efficient statistical evaluation of AI systems

Even when it is clear what a “good” output looks like, determining the “ground truth” performance of a system often requires time-intensive expert annotation. Our group develops statistical methods that make the most of limited labeled data, by e.g., carefully selecting a small number of expert labels in combination with LLM-as-judge or other automated measures of performance, while still providing valid statistical guarantees.

Fixed Size Active Statistical Inference
Erik Skalnes, Michael Oberst
Neural Information Processing Systems (NeurIPS), 2026

No Free Lunch: Non-Asymptotic Analysis of Prediction-Powered Inference
Pranav Mani, Peng Xu, Zachary C. Lipton, Michael Oberst
International Conference on Machine Learning (ICML), 2026
paper

Methods for causal evaluation of evolving AI systems

Even when an AI system scores well according to retrospective evaluations, it may fail to improve outcomes when deployed in the wild. Randomized evaluations (e.g., RCTs / AB testing) are an important tool for establishing these effects, but running randomized trials for every update to an AI system is not feasible. With that in mind, our group develops causal inference methods for assessing the real-world impact of AI systems, including their effect on downstream decisions and how to validate them on an ongoing basis after deployment.

Bounding the Causal Impact of ML-assisted Decision-Making via Counterfactual Correctness
Jonathan Zhang, Erik Skalnes, Jacob Chen, Michael Oberst
Conference on Uncertainty in Artificial Intelligence (UAI), 2026
paper

Just Trial Once: Ongoing Causal Validation of Machine Learning Models
Jacob M. Chen, Michael Oberst
Conference on Uncertainty in Artificial Intelligence (UAI), 2025
Oral Presentation (3% of submissions, 9% of accepted papers)
paperposter

Research Group

Meet the current members of the group and our collaborators on the People page. If you are interested in working with me as a PhD student or postdoc, please see this page for more information.