Predicting multi-cancer risk from EHR data using multi-agent LLMs.

S Sihang Zeng (Truveta Inc., Bellevue, WA) Y Youngwon Kim (School of Biological Sciences, Seoul National University) W Wilson Lau (Truveta Inc., Bellevue, WA) E Ehsan Alipour (Truveta Inc., Bellevue, WA) R Ruth Douglas Etzioni (University of Washington School of Public Health and Fred Hutchinson Cancer Center, Seattle, WA) M Meliha Yetisgen (University of Washington, Seattle, WA) A Anand Oka (Truveta Inc., Bellevue, WA) J Jay Nanduri (Truveta Inc., Bellevue, WA)

Abstract

10501 Background: Accurate multi-cancer risk prediction is essential for optimizing screening strategies and yet remains challenging due to heterogeneous and longitudinal electronic health record (EHR) data. Large language models (LLMs) offer a unified framework for multi-cancer risk assessment without disease-specific model training, but single-agent LLMs struggle to reason effectively over long patient histories. Methods: We developed a multi-agent system (MAS) composed of chained LLM agents with episodic memory to synthesize longitudinal EHR data into comprehensive cancer risk profiles. The system generates 1-year cancer risk scores (1–10) for individual cancer types. Using Truveta Data, we identified cases for 9 cancer types with clinically recognized precursor signs using clinician-curated diagnostic codes. For each cancer type, 500 cases were randomly sampled and matched 1:1 with controls by 10-year age group and sex. All structured EHR data (conditions, laboratory results, observations, medications, and procedures) up to 1 year before diagnosis or index date were included. Performance was evaluated across cancer types; lung cancer was used as a benchmark cohort for comparison with traditional machine learning (ML) models and a single-agent LLM baseline. Results: The MAS demonstrated heterogeneous but clinically meaningful discrimination across cancer types (AUROC range: 0.64 to 0.80), with liver cancer performing the best and colorectal cancer the worst. In the lung cancer cohort, performance (AUROC=0.79) was comparable to that of trained ML models (AUROC: XGBoost 0.82, logistic regression 0.73, k-nearest neighbors 0.58). Beyond risk prediction, the system generated concise patient summaries and interpretable rationales that supported downstream analyses, including cancer real-world evidence generation and exploratory knowledge discovery. Compared with a single-agent LLM, the MAS showed improved temporal coherence and clinical reasoning. Conclusions: A multi-agent LLM system can leverage longitudinal EHR data to estimate short-term multi-cancer risk within a unified framework, without cancer-specific model training. These findings support the potential role of LLM-based systems as scalable tools to inform risk stratification and screening strategies in real-world clinical populations. The MAS’s performance. Lung Ovarian Liver Pancreatic Colorectal Multiple Myeloma Lymphoma Gastric Bladder AUROC [95% CI] 0.79 [0.77, 0.82] 0.68 [0.65, 0.71] 0.80 [0.77, 0.83] 0.66 [0.62, 0.69] 0.64 [0.61, 0.68] 0.70 [0.67, 0.73] 0.69 [0.66, 0.72] 0.69 [0.66, 0.72] 0.71 [0.68, 0.74]

Article Details

Volume / Issue Vol. 44, Issue 16_suppl
Published June 01, 2026
Pages 10501-10501
ISSN 0732-183X
Publisher Lippincott Williams & Wilkins

Journal Info

Journal of Clinical Oncology

Lippincott Williams & Wilkins

ISSN: 0732-183X Health Sciences

Authors (8)

S

Sihang Zeng

Truveta Inc., Bellevue, WA

Y

Youngwon Kim

School of Biological Sciences, Seoul National University

W

Wilson Lau

Truveta Inc., Bellevue, WA

E

Ehsan Alipour

Truveta Inc., Bellevue, WA

R

Ruth Douglas Etzioni

University of Washington School of Public Health and Fred Hutchinson Cancer Center, Seattle, WA

M

Meliha Yetisgen

University of Washington, Seattle, WA

A

Anand Oka

Truveta Inc., Bellevue, WA

J

Jay Nanduri

Truveta Inc., Bellevue, WA