Enabling national identification of lung cancer screening eligibility with large language models.
Abstract
e13613 Background: Lung cancer (LC) is the leading cause of cancer-related death, yet screening uptake among eligible individuals is under 20%. Incomplete or inaccessible smoking history documentation, needed for determining eligibility, is a major barrier. Automating smoking history extraction in health records with natural language processing (NLP) could improve eligibility identification. However, traditional NLP systems struggle with generalizability across practice settings, limiting their effectiveness. Large language models (LLMs) show promise for overcoming these challenges in publicly available datasets, but their application to real-world, protected data remains untested, especially in large, diverse settings like the national Veterans Affairs (VA) system. Methods: LLM generalizability for smoking history extraction was assessed using protected clinical notes from over 100 VA sites (years 2008–2023) and the public non-VA MIMIC-III dataset (2001–2008). NLP tasks included status, duration, intensity, quit year, and pack-years. Using the open-source LLM Mixtral 8x22B, zero-shot prompts was compared to adapted prompts designed to address errors, measuring task-specific performance by F1. Institution generalizability was defined as the absolute per-task F1 score difference between the VA (training) and MIMIC (validation) datasets. To identify institutional factors influencing performance, we categorized errors using the concept of a hazard framework, distinguishing errors (or “hazards”) caused by LLM generalization limitations vs. those due to institution-specific differences. The effectiveness of identifying LC screening eligibility was compared to existing VA tools. Results: Adapted prompts improved generalizability across institutions, reducing the median F1 difference between VA and MIMIC datasets from 0.09 (interquartile range [IQR]: 0.08–0.14) with zero-shot prompts to 0.04 (IQR: 0.02–0.07). Overall accuracy with adapted prompts was similar across datasets (macro-F1: 0.86 in VA, 0.85 in MIMIC), but hazard distributions differed significantly (i.e., templated information was more prevalent in the VA dataset [p=0.004)]. Performance differed when task-specific requirements conflicted with pre-trained model behavior (p=0.007). Using LLM, we identified LC screening eligibility in 59% of Veterans who later developed LC, a 6.4-fold improvement over current VA tools, which identified only 8%. Conclusions: Adapted LLM prompts improve generalizability across diverse datasets, addressing longstanding challenges in clinical NLP. Using an open-source model deployable in protected healthcare settings, the robust LLM system maintained performance across years of data, allowing automated extraction of smoking history. This increased LC screening eligibility identification and can enable earlier detection to benefit over a million additional Veterans. The views expressed in this research study do not represent the view of the Department of Veterans Affairs or the United States Government.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (15)
Julie Tsu-yu Wu
Veterans Affairs Palo Alto Health Care System, Palo Alto, CA
Sydney Conover
VA Palo Alto Health Care System, Palo Alto, CA
Chloe Su
Stanford University School of Medicine, Palo Alto, CA
June Corrigan
2VA Boston Healthcare System, Boston, United States
John Culnan
2VA Boston Healthcare System, Boston, United States
Yuhan Liu
Michael J. Kelley
National Oncology Program Office, Department of Veterans Affairs, Durham VA Health Care System, Duke University, Durham, NC
Nhan Do
Massachusetts Veterans Epidemiology Research and Information Center, VA Boston Healthcare System, Boston, Massachusetts, United States
Shipra Arya
Alex HS Harris
VA Palo Alto Health Care System, Palo Alto, CA
Curtis Langlotz
Renda Soylemez Wiener
Boston University School of Medicine, Boston, MA
Westyn Branch-Elliman
UCLA, Los Angeles, CA
Summer Han
Stanford University School of Medicine, Stanford, CA
Nathanael Fillmore
Massachusetts Veterans Epidemiology Research and Information Center, VA Boston Healthcare System, Boston, Massachusetts, United States