Development and validation of predictive models for diabetic retinopathy using machine learning

P Penglu Yang B Bin Yang

Abstract

Objective This study aimed to develop and compare machine learning models for predicting diabetic retinopathy (DR) using clinical and biochemical data, specifically logistic regression, random forest, XGBoost, and neural networks. Methods A dataset of 3,000 diabetic patients, including 1,500 with DR, was obtained from the National Population Health Science Data Center. Significant predictors were identified, and four predictive models were developed. Model performance was assessed using accuracy, precision, recall, F1-score, and area under the curve (AUC). Results Random forest and XGBoost demonstrated superior performance, achieving accuracies of 95.67% and 94.67%, respectively, with AUC values of 0.991 and 0.989. Logistic regression yielded an accuracy of 76.50% (AUC: 0.828), while neural networks achieved 82.67% accuracy (AUC: 0.927). Key predictors included 24-hour urinary microalbumin, HbA1c, and serum creatinine. Conclusion The study highlights random forest and XGBoost as effective tools for early DR detection, emphasizing the importance of renal and glycemic markers in risk assessment. These findings support the integration of machine learning models into clinical decision-making for improved patient outcomes in diabetes management.

Article Details

Journal PLoS ONE
Volume / Issue Vol. 20, Issue 2
Published February 24, 2025
Pages e0318226
ISSN 1932-6203
Publisher Public Library of Science

Journal Info

PLoS ONE

Public Library of Science

ISSN: 1932-6203 Open Access Health Sciences

Authors (2)

P

Penglu Yang

B

Bin Yang