Enabling national identification of lung cancer screening eligibility with large language models.

J Julie Tsu-yu Wu (Veterans Affairs Palo Alto Health Care System, Palo Alto, CA) S Sydney Conover (VA Palo Alto Health Care System, Palo Alto, CA) C Chloe Su (Stanford University School of Medicine, Palo Alto, CA) J June Corrigan (2VA Boston Healthcare System, Boston, United States) J John Culnan (2VA Boston Healthcare System, Boston, United States) Y Yuhan Liu M Michael J. Kelley (National Oncology Program Office, Department of Veterans Affairs, Durham VA Health Care System, Duke University, Durham, NC) N Nhan Do (Massachusetts Veterans Epidemiology Research and Information Center, VA Boston Healthcare System, Boston, Massachusetts, United States) S Shipra Arya A Alex HS Harris (VA Palo Alto Health Care System, Palo Alto, CA) C Curtis Langlotz R Renda Soylemez Wiener (Boston University School of Medicine, Boston, MA) W Westyn Branch-Elliman (UCLA, Los Angeles, CA) S Summer Han (Stanford University School of Medicine, Stanford, CA) N Nathanael Fillmore (Massachusetts Veterans Epidemiology Research and Information Center, VA Boston Healthcare System, Boston, Massachusetts, United States)

Abstract

e13613 Background: Lung cancer (LC) is the leading cause of cancer-related death, yet screening uptake among eligible individuals is under 20%. Incomplete or inaccessible smoking history documentation, needed for determining eligibility, is a major barrier. Automating smoking history extraction in health records with natural language processing (NLP) could improve eligibility identification. However, traditional NLP systems struggle with generalizability across practice settings, limiting their effectiveness. Large language models (LLMs) show promise for overcoming these challenges in publicly available datasets, but their application to real-world, protected data remains untested, especially in large, diverse settings like the national Veterans Affairs (VA) system. Methods: LLM generalizability for smoking history extraction was assessed using protected clinical notes from over 100 VA sites (years 2008–2023) and the public non-VA MIMIC-III dataset (2001–2008). NLP tasks included status, duration, intensity, quit year, and pack-years. Using the open-source LLM Mixtral 8x22B, zero-shot prompts was compared to adapted prompts designed to address errors, measuring task-specific performance by F1. Institution generalizability was defined as the absolute per-task F1 score difference between the VA (training) and MIMIC (validation) datasets. To identify institutional factors influencing performance, we categorized errors using the concept of a hazard framework, distinguishing errors (or “hazards”) caused by LLM generalization limitations vs. those due to institution-specific differences. The effectiveness of identifying LC screening eligibility was compared to existing VA tools. Results: Adapted prompts improved generalizability across institutions, reducing the median F1 difference between VA and MIMIC datasets from 0.09 (interquartile range [IQR]: 0.08–0.14) with zero-shot prompts to 0.04 (IQR: 0.02–0.07). Overall accuracy with adapted prompts was similar across datasets (macro-F1: 0.86 in VA, 0.85 in MIMIC), but hazard distributions differed significantly (i.e., templated information was more prevalent in the VA dataset [p=0.004)]. Performance differed when task-specific requirements conflicted with pre-trained model behavior (p=0.007). Using LLM, we identified LC screening eligibility in 59% of Veterans who later developed LC, a 6.4-fold improvement over current VA tools, which identified only 8%. Conclusions: Adapted LLM prompts improve generalizability across diverse datasets, addressing longstanding challenges in clinical NLP. Using an open-source model deployable in protected healthcare settings, the robust LLM system maintained performance across years of data, allowing automated extraction of smoking history. This increased LC screening eligibility identification and can enable earlier detection to benefit over a million additional Veterans. The views expressed in this research study do not represent the view of the Department of Veterans Affairs or the United States Government.

Article Details

Volume / Issue Vol. 43, Issue 16_suppl
Published June 01, 2025
ISSN 0732-183X
Publisher Lippincott Williams & Wilkins

Journal Info

Journal of Clinical Oncology

Lippincott Williams & Wilkins

ISSN: 0732-183X Health Sciences

Authors (15)

J

Julie Tsu-yu Wu

Veterans Affairs Palo Alto Health Care System, Palo Alto, CA

S

Sydney Conover

VA Palo Alto Health Care System, Palo Alto, CA

C

Chloe Su

Stanford University School of Medicine, Palo Alto, CA

J

June Corrigan

2VA Boston Healthcare System, Boston, United States

J

John Culnan

2VA Boston Healthcare System, Boston, United States

Y

Yuhan Liu

M

Michael J. Kelley

National Oncology Program Office, Department of Veterans Affairs, Durham VA Health Care System, Duke University, Durham, NC

N

Nhan Do

Massachusetts Veterans Epidemiology Research and Information Center, VA Boston Healthcare System, Boston, Massachusetts, United States

S

Shipra Arya

A

Alex HS Harris

VA Palo Alto Health Care System, Palo Alto, CA

C

Curtis Langlotz

R

Renda Soylemez Wiener

Boston University School of Medicine, Boston, MA

W

Westyn Branch-Elliman

UCLA, Los Angeles, CA

S

Summer Han

Stanford University School of Medicine, Stanford, CA

N

Nathanael Fillmore

Massachusetts Veterans Epidemiology Research and Information Center, VA Boston Healthcare System, Boston, Massachusetts, United States