Understanding the role of large language models in cancer mortality prediction: A real-world study.
Abstract
e23177 Background: Although mortality risk prediction is essential in oncology, risk determinants vary widely across cancer types and patient characteristics, rendering population-level prediction challenging. While traditional statistical and machine learning (ML) approaches have shown encouraging results, the effectiveness of ML and large language models (LLMs) for near and long-term cancer mortality prediction remains unknown. We sought to assess the mortality prediction of these different models for multiple timepoints in a heterogeneous cancer population. Methods: We conducted a retrospective cohort study using electronic health record (EHR) data from Yale New Haven Health, with follow-up through 6/30/25. Patients with any first cancer diagnosis on or after 1/2019 were included. One post-diagnosis encounter was randomly selected as the index visit, excluding visits within 7 days of death. The final cohort was randomly divided into training, validation, and test sets. Structed demographics, laboratory, and diagnostic comorbidity codes were included in training models and derived from the 180 days preceding the index visit. Outcomes included mortality at 30 and 180 days after index visit. We assessed logistic regression (LR), ML model (XGBoost), and LLM (GPT-4o) for prediction of mortality. LR and XGBoost were trained on the full training set, whereas GPT-4o was not trained on patient data and evaluated using zero-shot inference and limited in-context learning (ICL) with five example patients, within a secure environment. Model performance was compared using area under the receiver operating characteristic curve (AUROC), sensitivity, and specificity with higher values indicating better performance. Results: Our cohort included 82,662 individuals with a median follow-up of 23.4 months after their index date. 15,152 (18.33%) of individuals died during our study period. The most common diagnoses were prostate (14.5%), breast (12.4%), and lung (6.8%). Fully trained ML models achieved the highest AUROC, while untrained LLMs (0-shot) showed meaningful performance and improved with minimal ICL, approaching traditional and ML methods (Table). Conclusions: Our findings highlight the importance of assessing both short- and longer-term mortality, as model performance differed across timepoints. LLMs captured clinically relevant risk signals from structured EHR data even without task-specific training, with further improvement using minimal ICL. While ML achieved the best performance after supervised training, LLMs show promise for mortality prediction and integration into EHR-based clinical decision support. Performance of models for mortality prediction. Model Time (days) AUROC Sensitivity Specificity LR 30 0.90 0.81 0.84 180 0.87 0.74 0.82 XGBoost 30 0.92 0.38 0.98 180 0.89 0.59 0.93 GPT-4o (0-shot) 30 0.87 0.90 0.70 180 0.83 0.81 0.73 GPT-4o (ICL) 30 0.88 0.85 0.80 180 0.84 0.75 0.80
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (8)
Xueqing Peng
Maureen Canavan
Yale School of Medicine, New Haven, CT
Huan He
National Engineering Laboratory for Druggable Gene and Protein Screening, College of Life Science, Northeast Normal University
Sarah Westvold
Yale Cancer Outcomes, Public Policy and Effectiveness Research Center, New Haven, CT
Lingfei Qian
Cary Philip Gross
National Clinician Scholars Program; Yale Cancer Outcomes, Public Policy and Effectiveness Research Center; Yale School of Medicine, New Haven, CT
Scott F. Huntington
Yale University, New Haven, CT
Hua Xu
State Key Laboratory of Gene Function and Modulation Research, School of Life Sciences, and Biomedical Pioneering Innovation Center, Peking University