Development of an oncology generative AI foundation model trained on more than a million longitudinal patient journeys across the United States.

W Wilson Lau (Truveta Inc., Bellevue, WA) E Ehsan Alipour (Truveta Inc., Bellevue, WA) Y Youngwon Kim (School of Biological Sciences, Seoul National University) S Sihang Zeng (Truveta Inc., Bellevue, WA) A Anand Oka (Truveta Inc., Bellevue, WA) J Jay Nanduri (Truveta Inc., Bellevue, WA)

Abstract

10558 Background: Incidence of cancer diagnoses continues to increase, while the average cost of cancer screening, such as mammogram or colonoscopy, remains high, ranging from hundreds to over a thousand dollars. This study explores the potential of building an oncology foundation model based on Generative Pre-trained Transformers (GPT) to predict future outcomes for patients. We assess the prediction accuracy and feasibility of leveraging the foundation model to inform when screening should be prioritized, thereby reducing the associated burden of unnecessary procedures. Methods: The advancement of generative AI offers the opportunity to model the progression of human disease over time. In this study, we extended the GPT architecture and pre-trained it with a subset of Truveta Data containing the electronic health record (EHR) from the journeys of 1.4 million de-identified patients diagnosed with 4 types of cancer (lung, breast, colorectal, prostate). Each sequence of the patient journey—including demographics (age, gender), conditions, lab results (represented by SNOMED-CT and LOINC codes), and the corresponding event time—was tokenized before passing through the GPT model for training. For validation, 1000 patients were randomly sampled from our test data, in which each type of cancer diagnosis constituted 19%-30% of the samples. We used the model to generate synthetic future patient journeys for the selected patients and compared the predicted outcomes with the actual cancer diagnoses within one year. Results: The model achieved sensitivities above 78%, with positive predictive values (PPV) between 66% and 90%. More importantly, it demonstrated high specificities above 91%, with corresponding negative predictive values (NPV) above 90%. Conclusions: The high specificities and NPV indicate the feasibility of applying generative foundation model pre-trained with EHR data to predict negative cancer outcomes with high accuracy. Since the percentage of screening tests leading to positive cancer diagnosis is relatively low, the projected negative predictive outcomes offer valuable signals for clinicians, which they can use to complement their expert assessment to avoid unnecessary screening and subsequently reduce the burden of screening costs. Performance of the foundation model on cancer outcome prediction. Lung Breast Colorectal Prostate Positive predictive value (PPV) 72.77% 90.11% 65.78% 88.88% Sensitivity 80.28% 79.00% 78.13% 78.26% Negative predictive value (NPV) 91.87% 96.29% 90.35% 96.27% Specificity 94.50% 91.45% 94.56% 92.07%

Article Details

Volume / Issue Vol. 44, Issue 16_suppl
Published June 01, 2026
Pages 10558-10558
ISSN 0732-183X
Publisher Lippincott Williams & Wilkins

Journal Info

Journal of Clinical Oncology

Lippincott Williams & Wilkins

ISSN: 0732-183X Health Sciences

Authors (6)

W

Wilson Lau

Truveta Inc., Bellevue, WA

E

Ehsan Alipour

Truveta Inc., Bellevue, WA

Y

Youngwon Kim

School of Biological Sciences, Seoul National University

S

Sihang Zeng

Truveta Inc., Bellevue, WA

A

Anand Oka

Truveta Inc., Bellevue, WA

J

Jay Nanduri

Truveta Inc., Bellevue, WA