Data subdivision approach enhances machine learning-based mortality prediction in pediatric ICU patients
Abstract
Objective To evaluate machine learning–based models for predicting all-cause mortality in pediatric ICU patients using comprehensive biochemical panels, with a focus on addressing missing data and class imbalance. Materials and methods A retrospective analysis was performed on a publicly available PICU dataset comprising 8,629 patients aged 28 days to 18 years. Twenty-two biochemical variables measured on the first PICU day were analyzed. Missing values were addressed using Multiple Imputation. Lasso regression was applied for feature selection. Models were trained using 5-fold cross-validation with the Synthetic Minority Oversampling Technique (SMOTE). Two ML-based imbalance-handling strategies, stacking ensemble and data subdivision were evaluated. Pairwise DeLong tests were used to compare AUC performance across models. Results Among the 8,629 included patients, there were 476 non-survivors (5.5 percent). Multiple imputation followed by SMOTE improved model performance across all algorithms. The single-model classifiers achieved AUC-ROC values of 0.82 (Random Forest), 0.79 (CatBoost), 0.83 (Extra Trees), and 0.79 (Logistic Regression). The stacking ensemble demonstrated the best overall performance, with an AUC-ROC of 0.88 and an AUC-PRC of 0.45. The data subdivision approaches also produced strong discriminative performance, achieving AUC-ROC value up to 0.83 for three-subdivision, and 0.82 for five-subdivision strategies. Calibration analysis showed that the stacking model achieved the lowest Brier score (0.04), indicating superior probabilistic accuracy compared with individual classifiers. Feature importance analyses across all MI-based models consistently highlighted coagulation markers (D-dimer, reference TT, PTT, INR), electrolytes (chloride, potassium, sodium), and metabolic and organ-dysfunction indicators (AST, ALT, creatinine) as key predictors of mortality. Conclusions This study demonstrates that ensemble stacking is a more effective strategy than data subdivision for addressing class imbalance in PICU mortality prediction.
Article Details
Authors (7)
Wenqian Chen
Benjamin Lee
Zexi Zang
Junfeng Li
Tsinghua Shenzhen International Graduate School
Lingna Huang
Hang Xing
State Key Laboratory of Chemo-/Bio-Sensing and Chemometrics, School of Chemistry and Chemical Engineering
Yanli Ren