Predicting multi-cancer risk from EHR data using multi-agent LLMs.
Abstract
10501 Background: Accurate multi-cancer risk prediction is essential for optimizing screening strategies and yet remains challenging due to heterogeneous and longitudinal electronic health record (EHR) data. Large language models (LLMs) offer a unified framework for multi-cancer risk assessment without disease-specific model training, but single-agent LLMs struggle to reason effectively over long patient histories. Methods: We developed a multi-agent system (MAS) composed of chained LLM agents with episodic memory to synthesize longitudinal EHR data into comprehensive cancer risk profiles. The system generates 1-year cancer risk scores (1–10) for individual cancer types. Using Truveta Data, we identified cases for 9 cancer types with clinically recognized precursor signs using clinician-curated diagnostic codes. For each cancer type, 500 cases were randomly sampled and matched 1:1 with controls by 10-year age group and sex. All structured EHR data (conditions, laboratory results, observations, medications, and procedures) up to 1 year before diagnosis or index date were included. Performance was evaluated across cancer types; lung cancer was used as a benchmark cohort for comparison with traditional machine learning (ML) models and a single-agent LLM baseline. Results: The MAS demonstrated heterogeneous but clinically meaningful discrimination across cancer types (AUROC range: 0.64 to 0.80), with liver cancer performing the best and colorectal cancer the worst. In the lung cancer cohort, performance (AUROC=0.79) was comparable to that of trained ML models (AUROC: XGBoost 0.82, logistic regression 0.73, k-nearest neighbors 0.58). Beyond risk prediction, the system generated concise patient summaries and interpretable rationales that supported downstream analyses, including cancer real-world evidence generation and exploratory knowledge discovery. Compared with a single-agent LLM, the MAS showed improved temporal coherence and clinical reasoning. Conclusions: A multi-agent LLM system can leverage longitudinal EHR data to estimate short-term multi-cancer risk within a unified framework, without cancer-specific model training. These findings support the potential role of LLM-based systems as scalable tools to inform risk stratification and screening strategies in real-world clinical populations. The MAS’s performance. Lung Ovarian Liver Pancreatic Colorectal Multiple Myeloma Lymphoma Gastric Bladder AUROC [95% CI] 0.79 [0.77, 0.82] 0.68 [0.65, 0.71] 0.80 [0.77, 0.83] 0.66 [0.62, 0.69] 0.64 [0.61, 0.68] 0.70 [0.67, 0.73] 0.69 [0.66, 0.72] 0.69 [0.66, 0.72] 0.71 [0.68, 0.74]
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (8)
Sihang Zeng
Truveta Inc., Bellevue, WA
Youngwon Kim
School of Biological Sciences, Seoul National University
Wilson Lau
Truveta Inc., Bellevue, WA
Ehsan Alipour
Truveta Inc., Bellevue, WA
Ruth Douglas Etzioni
University of Washington School of Public Health and Fred Hutchinson Cancer Center, Seattle, WA
Meliha Yetisgen
University of Washington, Seattle, WA
Anand Oka
Truveta Inc., Bellevue, WA
Jay Nanduri
Truveta Inc., Bellevue, WA