Abstract 4366171: One Model Does Not Fit All: Demographic Disparities in Machine Learning-Based Prediction of 30-Day Readmission Following Acute Myocardial Infarction

J Jaini Shah (Dartmouth College, Monroe, New York, United States) M Michael Matheny R Ruth Reeves (Vanderbilt University Medical Center, Nashville, Tennessee, United States) J Jeremiah Brown (THE DARTMOUTH INSTITUTE, Lebanon, New Hampshire, United States) I Iben Ricket (Dartmouth College, Monroe, New York, United States)

Abstract

Background: Machine learning (ML) models are increasingly used to predict clinical outcomes such as 30-day readmission following acute myocardial infarction (AMI). However, most models are trained on heterogeneous populations and rarely account for demographic subgroup disparities potentially exacerbating inequities in care. This lack of subgroup-specific considerations may result in biased predictions, especially if a model underperforms in certain populations, leading to misinformed clinical decisions and widening existing health disparities. Objective: To evaluate the performance of a generalized XGBoost model for predicting 30-day AMI readmission and assess how accuracy varies across demographic and clinical subgroups. Methods: This study utilized a cohort of electronic health records from Vanderbilt University Medical Center. Data included patients hospitalized between 2007–2016 with an acute myocardial infarction and included variables on demographics (N=3) and clinical characteristics (N=108). The outcome was 30-day readmission from incident hospitalization. Missing variables were imputed with model-based imputation using K-Nearest Neighbors. After standard preprocessing and stratified train-test splitting (80/20), we trained an XGBoost classifier and evaluated performance using AUROC, specificity, and sensitivity. We then stratified the data by age group, sex, race, comorbidity category (Charlson), and length of stay to examine subgroup-specific model performance. Results: The cohort included 6,179 patients and 10.5% of them had a 30-day readmission. The overall XGBoost model achieved high accuracy (89.5%) and specificity (99.5%) but demonstrated poor sensitivity (3.8%) and a modest AUROC of 0.6050. Subgroup analysis showed major performance gaps. Subgroup analysis identified major performance differences, especially for younger patients (Age <40), older patients (Age>80), patients with greater co-morbidities (Charlson severe) and patients with a medium length of stay (LOS). Based on AUROC, the following sub-groups performed better than the overall model: Females, Age 40-60, Age 60-80, Charlson Mild, LOS shot, and LOS Long. Conclusions: Material differences in performance metrics were seen across demographic and clinical subgroups. This highlights the risk of deploying generalized models in diverse populations and underscores the need for subgroup-specific or fairness-aware modeling to improve both accuracy and equity in clinical predictions.

Article Details

Journal Circulation
Volume / Issue Vol. 152, Issue Suppl_3
Published November 04, 2025
ISSN 0009-7322
Publisher Lippincott Williams & Wilkins

Journal Info

Circulation

Lippincott Williams & Wilkins

ISSN: 0009-7322 Health Sciences

Authors (5)

J

Jaini Shah

Dartmouth College, Monroe, New York, United States

M

Michael Matheny

R

Ruth Reeves

Vanderbilt University Medical Center, Nashville, Tennessee, United States

J

Jeremiah Brown

THE DARTMOUTH INSTITUTE, Lebanon, New Hampshire, United States

I

Iben Ricket

Dartmouth College, Monroe, New York, United States