Automated eligibility screening for adjuvant CDK4/6 inhibitors in high-risk HR+/HER2− early breast cancer using natural language processing.
Abstract
e23303 Background: While CDK4/6 inhibitors (CDK4/6i) have transformed the treatment landscape for high-risk, HR+/HER2- early breast cancer (EBC), identifying eligible patients in real-world settings remains labor-intensive. Manual review of unstructured clinical notes to match patients with evolving FDA criteria (e.g., abemaciclib, March 2023; ribociclib, September 2024) has led to an eligibility screening gap. We evaluated an automated pipeline to extract clinical eligibility criteria and facilitate rapid patient-treatment matching. Methods: We developed a hybrid Python-based information extraction pipeline that integrates spaCy-based named entity recognition (NER) with customized rule- and regex-based algorithms to parse both structured electronic health record (EHR) fields and unstructured clinical narratives, including pathology reports, from Yale Cancer Center. The system extracted 11 clinically relevant variables (e.g., ER/PR/HER2 status, tumor stage, Ki-67 index, nodal involvement). Eligibility for adjuvant CDK4/6 inhibitor therapy was determined using rule-based logic aligned with U.S. FDA-approved indications for abemaciclib and ribociclib. A gold-standard manual chart review was conducted by a trained clinical data coordinator on 100 randomly selected early breast cancer (EBC) patients. Extraction accuracy, eligibility classification performance (eligible vs. ineligible), and processing time were evaluated by comparing automated outputs against the manual reference standard. Results: The pipeline achieved ≥ 90% accuracy across all 11 individual data points. For the final eligibility classification, the tool demonstrated 99% accuracy, with only one false negative (classified as ineligible by NLP but eligible by manual review). Efficiency gains were significant: manual chart review required a median of 3 minutes per patient (range: 1–12), whereas the automated pipeline processed 100 patients in ~1 minute, representing a > 99% reduction in screening time. Conclusions: Automated NLP extraction is an accurate and scalable solution for identifying patients eligible for adjuvant CDK4/6i. By substantially reducing the time required for chart review, this tool can minimize treatment delays and ensure real-world adherence to changing FDA guidelines. Future work will compare this rules-based approach with ML and LLM-based extractors to improve semantic context handling and scalability across health systems. Eligibility classification performance of regex-based pipeline compared to manual chart review. Metric NLP/Regex-based pipeline performance Accuracy 99.0% Sensitivity (Recall) 100.0% Specificity 98.9% Precision 91.7% F1 Score 95.7%
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (12)
Jessica Liu
Division of Medical Biology, Genomics Research Center, Academia Sinica
Sameer Pandya
Britny Brown
University of Rhode Island College of Pharmacy, Kingston, RI
Michelle Caetano
University of Rhode Island College of Pharmacy, Kingston, RI
Salma Taghzout
University of Rhode Island College of Pharmacy, Kingston, RI
Mariah Ramos
University of Rhode Island College of Pharmacy, Kingston, RI
Annette Hood
Yale New Haven Hospital, New Haven, CT
Michael Zummo
Yale New Haven Health, New Haven, CT
Wei Wei
Robert Duffy Legare
Yale School of Medicine, Westerly, RI
Maryam B. Lustberg
Yale Cancer Center, Yale School of Medicine, New Haven, CT
Guannan Gong