Machine learning–integrated routine biomarkers for breast cancer risk stratification: A scalable strategy for resource-limited settings.
Abstract
e12557 Background: Early breast cancer diagnosis is essential for improving survival outcomes and optimizing healthcare resource utilization. However, in resource-limited settings and low- and middle-income countries, access to screening programs and advanced risk assessment tools remains limited, and scalable strategies leveraging routinely collected data are underexplored. We evaluated routine blood-based biomarkers, individually and in combination, using a machine learning (ML) approach to develop a low-cost, scalable breast cancer risk assessment tool. Methods: This cross-sectional study included 236 women aged 40–70 years evaluated at Barretos Cancer Hospital between January 2024 and January 2025. Participants comprised newly diagnosed breast cancer patients (n = 125) and women undergoing screening with benign imaging findings (BI-RADS 1–2; n = 111). Routine laboratory markers included complete blood count, metabolic and inflammatory markers, thyroid and sex hormones, and tumor markers. Group comparisons were performed using Student’s t -test. A ridge regression model was trained using biomarkers significantly associated with breast cancer (p < 0.05). Results: The median age was 52 years. The majority of participants were White (51.7%), followed by Brown (37.1%) and Black (10.3%); 56.8% were postmenopausal. Several biomarkers were significantly higher in women with breast cancer (p < 0.05), including CA 15-3 (p = 0.018), CEA (p = 0.014), FSH (p = 0.016), T3 (p < 0.001) and Mean Corpuscular Hemoglobin Concentration - MCHC (p = 0.014). No strong correlations were observed among selected variables (maximum correlation 0.19 between CEA and CA 15-3). Incorporation of biomarkers with p < 0.05 (CA 15-3, CEA, FSH, T3 and MCHC) into the ridge regression model yielded a mean AUC of 0.72 (95% CI, 0.68–0.75) based on 1,000 iterations of stratified random subsampling, with sensitivity of 0.63 ± 0.08, specificity of 0.67 ± 0.08, accuracy of 0.65 ± 0.05, balanced accuracy of 0.65 ± 0.05, negative predictive value of 0.63 ± 0.05, and positive predictive value of 0.67 ± 0.06. Conclusions: An ML model combining routine blood-based biomarkers demonstrated potential as a low-cost pre-screening tool for breast cancer risk stratification. This approach may help prioritize diagnostic resources, particularly in resource-limited settings where access to advanced imaging and invasive diagnostic procedures is constrained. External validation in independent and diverse populations is required to confirm generalizability and support future clinical application.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (11)
Diego Renan Silva
HUNA, São Paulo, São Paulo, Brazil
Caroline Rogeri
HUNA, São Paulo, São Paulo, Brazil
Suzylaine da Silva Lima
HUNA, São Paulo, São Paulo, Brazil
Vinicius Moura Ribeiro
HUNA, São Paulo, São Paulo, Brazil
Marco Aurelio Kohara
HUNA, São Paulo, São Paulo, Brazil
Thiago Silva
Barretos Cancer Hospital, Barretos, Brazil
Vinicius L. Vazquez
Barretos Cancer Hospital, Barretos, Brazil
Guilherme Hernandes Garcia Sanchez
Barretos Cancer Hospital, Barretos, Brazil
Karina Braga Gomes
Federal University of Minas Gerais (UFMG), Belo Horizonte, Minas Gerais, Brazil
Pedro Henrique Araujo De Souza
Brazilian National Cancer Institute (INCA), Rio De Janeiro, Rio de Janeiro, Brazil
Daniella Araújo
HUNA, São Paulo, São Paulo, Brazil