Abstract MPTU06: Prediction of cardiovascular diseases using statistical and machine learning approaches: The Strong Heart Study

M Mohammad Anwarul Islam (University of Oklahoma Health Campus, Oklahoma City, Oklahoma, United States) S Steven Pan (Northwestern Medicine, Chicago, Illinois, United States) P Paul Rogers T Tauqeer Ali S Shelley Cole (Texas Biomedical Research Institute, San Antonio, Texas, United States) A Amanda Fretts (University of Washington School of Public Health, Seattle, Washington, United States) J Jessica Reese (University of Oklahoma- HSC, Oklahoma City, Oklahoma, United States) J Jason Umans (MedStar Health Research Institute, Bethesda, Maryland, United States) Y Ying Zhang

Abstract

Background: Machine learning (ML) models are highly non-linear functions that are known to have good prediction ability in the classification paradigms, from binary to multi-class classification of a response variable. Predicting cardiovascular disease (CVD) outcomes based on demographic and clinical characteristics is important but also challenging. ML and deep learning (DL) models can also be sensitive to the population structure of the dataset. In recent years, though ML approaches have been used in CVD prediction and its associated risk factor analysis for various populations, this technology could benefit underserved and rural populations. Therefore, the aim of this study was to use ML and DL models to predict CVD outcomes of American Indians in the Strong Heart Study (SHS), a longitudinal study of CVD conducted in Arizona, Oklahoma, North Dakota, and South Dakota. Methods: A total of 3,248 SHS participants were initially examined between 1989 and 1991. They were followed through 2023. Multivariable logistic regression was used to identify important features that are associated with CVD outcomes, in addition to literature review. These selected features were then used to train and test the ML and DL models, which were applied to predict CVD, Coronary Heart Disease (CHD), and stroke events. Sensitivity, specificity, F1 scores, and ROC curves were examined to evaluate the prediction accuracy of these models. Results: Among ML models, the support vector machine (SVM) performed best with the highest accuracy of 64% and 67% in predicting CVD and CHD outcomes, respectively, which are similar to the accuracy of logistic regression models. Among DL models, artificial neural networks (ANN) achieved the highest accuracy of 63% in predicting CVD, while convolutional neural networks (CNN) performed best for CHD with 65% accuracy. Both ML and DL models performed similarly in predicting stroke, with an average accuracy of 89.41%. Conclusion: Our results showed that ML and DL prediction models performed better when the disease prevalence was low. ML and DL models provide additional tools for predicting rare disease outcomes.

Article Details

Journal Circulation
Volume / Issue Vol. 153, Issue Suppl_1
Published March 24, 2026
ISSN 0009-7322
Publisher Lippincott Williams & Wilkins

Journal Info

Circulation

Lippincott Williams & Wilkins

ISSN: 0009-7322 Health Sciences

Authors (9)

M

Mohammad Anwarul Islam

University of Oklahoma Health Campus, Oklahoma City, Oklahoma, United States

S

Steven Pan

Northwestern Medicine, Chicago, Illinois, United States

P

Paul Rogers

T

Tauqeer Ali

S

Shelley Cole

Texas Biomedical Research Institute, San Antonio, Texas, United States

A

Amanda Fretts

University of Washington School of Public Health, Seattle, Washington, United States

J

Jessica Reese

University of Oklahoma- HSC, Oklahoma City, Oklahoma, United States

J

Jason Umans

MedStar Health Research Institute, Bethesda, Maryland, United States

Y

Ying Zhang