Machine learning for cancer risk stratification: A bi-directional approach to screening.
Abstract
10526 Background: Cancer screening programs are often limited by high costs, invasiveness, and potential false results. Machine learning (ML) applied to routine medical data could enable risk-stratified screening approaches. We developed ML models to predict 10-year cancer risk using data from routine periodic health examinations, aiming to identify both high and low-risk populations. Methods: We analyzed data from individuals who underwent routine health examinations at Rambam Medical Center (2002-2021), matched with the Israeli National Cancer Registry. After quality control, including removal of cases with previous tumors, we excluded cancers diagnosed less than 183 days after the visit to avoid prevalent cases. Using data from the initial visit only, we developed two XGBoost models for 10-year cancer risk: Model 1 (baseline) using only age and gender, and Model 2 incorporating 53 features including demographics, lifestyle factors (smoking, alcohol, physical activity), laboratory results (26 parameters including complete blood count, biochemistry, and lipids), medication categories (9 groups including statins, antihypertensives, antidiabetics), and medical history (7 disease categories). Models were trained on 75% of the data with cross-validation and evaluated on a 25% hold-out set. Results: Our cohort included 27,901 individuals (61% male, mean age 47±10 years), with 1,960 future incidents of cancer diagnosed at median follow-up of 8 years. Most common malignancies were prostate (21%), breast (19%), skin (15%), and colorectal (12%) cancers. For model development, we reduced the dataset to include only individuals with 10-year follow-up or cancer event within the 10-year window (N = 16,859, cancer cases = 1,268). Model 2 significantly outperformed Model 1 (AUC 0.799 vs 0.706, p < 0.0001) and demonstrated remarkable risk stratification. Against a population baseline risk of 7.3%, Model 2 identified: 1) a low-risk group (bottom 50%) with only 1.9% 10-year cancer risk; 2) an high risk group (90-99th percentile) with 25.1% risk (3.5-fold increase); and 3) a very high-risk group (top 1%) with 74.4% risk (10-fold increase). In comparison, the highest-risk group in Model 1 (top 1%) achieved only 20.0% risk. The most important predictive features were age, monocyte count, albumin, total bilirubin, and LDL cholesterol. Conclusions: ML analysis of routine health examination data can effectively stratify cancer risk, identifying both very low and exceptionally high-risk groups. This bi-directional stratification could enable more efficient screening strategies: reduce screening intervals or invasiveness in the large low-risk population, while intensifying screening in high-risk groups. Future studies should evaluate and validate these findings in different populations and study whether this approach can improve the efficiency and cost-effectiveness of cancer screening programs.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (4)
Erez Hasnis
Rambam Health Care Campus, Haifa, Israel
Ophir Avizohar
Rambam Health Care Campus, Haifa, Israel
Elizabeth E. Half
Rambam Health Care Campus, Haifa, Israel
Dvir Aran