Using machine learning to predict depression among middle-aged and elderly population in China and conducting empirical analysis

Z Zhe Wang N Ni Jia

Abstract

Objective To develop a predictive model for evaluating depression among middle-aged and elderly individuals in China. Methods Participants aged ≥ 45 from the 2020 China Health and Retirement Survey (CHARLS) cross-sectional study were enrolled. Depressive mood was defined as a score of 10 or higher on the CESD-10 scale, which has a maximum score of 30. A predictive model was developed using five selected machine learning algorithms. The model was trained and validated on the 2020 database cohort and externally validated through a questionnaire survey of middle-aged and elderly individuals in Shaanxi Province, China, following the same criteria. SHapley Additive Interpretation (SHAP) was employed to assess the importance of predictive factors. Results The stacked ensemble model demonstrated an AUC of 0.8021 in the test set of the training cohort for predicting depressive symptoms; the corresponding AUC in the external validation cohort was 0.7448, outperforming all base models. Conclusion The stacked ensemble approach serves as an effective tool for identifying depression in a large population of middle-aged and elderly individuals in China. For depression prediction, factors such as life satisfaction, self-reported health, pain, sleep duration, and cognitive function are identified as highly significant predictive factors.

Article Details

Journal PLoS ONE
Volume / Issue Vol. 20, Issue 3
Published March 18, 2025
Pages e0319232
ISSN 1932-6203
Publisher Public Library of Science

Journal Info

PLoS ONE

Public Library of Science

ISSN: 1932-6203 Open Access Health Sciences

Authors (2)

Z

Zhe Wang

N

Ni Jia