Stability and interpretability of penalized logistic regression models for breast cancer risk prediction
Abstract
Penalized logistic regression is widely used in biomedical classification to address multicollinearity and improve predictive performance, yet the stability and reproducibility of selected predictors are often overlooked. This study evaluates feature stability and interpretability in ridge, lasso, and elastic-net logistic regression for breast cancer diagnosis using the Wisconsin Diagnostic Breast Cancer dataset. Models were trained with cross-validated tuning and evaluated on an independent test set using discrimination, classification, and calibration metrics. Feature stability was quantified through bootstrap selection frequencies. All penalized models achieved near-perfect discrimination and improved calibration compared with unpenalized logistic regression. However, substantial differences emerged in stability and sparsity. Ridge regression exhibited maximal stability but retained all predictors, limiting interpretability. Lasso regression produced highly sparse models but showed greater selection variability. Elastic-net regression balanced sparsity and stability, consistently retaining correlated predictors linked to tumor morphology. These findings demonstrate that stability assessment provides critical information beyond predictive accuracy and supports stability-aware penalized modeling for interpretable and reproducible biomedical risk prediction.
Article Details
Authors (2)
Francis Okyere
Michael Nyanney