Impact of consensus-based feature selection on machine learning performance for metastatic lung cancer prediction in resource-limited settings.
Abstract
e20519 Background: In low- and middle-income countries (LMICs), lung cancer is frequently diagnosed at advanced or metastatic stages due to limited early detection infrastructure. Conventional logistic regression provides interpretability but limited predictive performance in high-dimensional settings, whereas penalized regression and machine-learning models can capture complex, non-linear relationships. Comparative evaluation of these approaches in LMIC metastatic cohorts remains scarce. We assessed regression-based and machine-learning models to predict metastatic lung cancer risk and identify key determinants, emphasizing the trade-off between predictive accuracy and clinical interpretability. Methods: We analyzed 4,480 lung cancer patients treated at the National Institute of Cancer Research and Hospital, Bangladesh (2021–2023) using 29 demographic, clinical, symptom-related, and laboratory variables. Missing binary and categorical data were imputed using observed-proportion sampling. Models evaluated included conventional logistic regression, LASSO-penalized logistic regression, ridge regression, Random Forest, and XGBoost. Feature importance was derived for each model and aggregated into a consensus ranking. The top 15 consensus-ranked features were used to re-train all models, with performance compared to full-feature models. Evaluation metrics included ROC–AUC, F1-score, sensitivity (metastatic recall), specificity, calibration plots, and decision curve analysis. Results: Among the top 15 consensus-ranked features, habitual factors emerged as the strongest predictors of metastatic disease, followed by poor performance status (ECOG ≥2), respiratory symptoms including breathlessness with chest pain, and multiple concurrent symptoms. After consensus-based feature selection, tree-based models demonstrated notable performance gains. Random Forest achieved a 3.6% increase in F1-score and 1.8% improvement in metastatic recall. XGBoost showed the largest relative gains: metastatic recall increased by 15.8%, F1-score by 8.3%, with modest improvements in discrimination. In contrast, regression models showed declines in F1-score (LASSO −7.9%, ridge −41.4%) despite stable or slightly improved specificity. Consensus-ranked features preferentially enhanced clinically relevant detection in tree-based models, demonstrating their superiority over regression approaches in this LMIC real-world dataset. Conclusions: Tree-based models achieved superior predictive accuracy and balanced sensitivity–specificity compared with conventional regression. Consensus-ranked features improved metastatic case detection despite dataset limitations inherent to LMIC settings. External validation is required to confirm generalizability in LMIC, real-world clinical environments.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (7)
Muhammad Rafiqul Islam
National Institute of Cancer Research and Hospital, Dhaka, Bangladesh
Syeda Masuma Siddiqua
Unity Through Population Service, Dhaka, Bangladesh
Mohammad Hasan Shahriar
University of Chicago, Chicago, IL
Muhammad Ashique Haider Chowdhury
University of Chicago, Chicago, IL
Siara Tasmin
University of Chicago, Chicago, IL
Humayera Islam
University of Chicago, Chicago, IL
Habibul Ahsan