Development and validation of machine learning risk prediction models for detection of early-onset colorectal cancer: Data from 30 health systems in the United States.
Abstract
3613 Background: Incidence of early-onset colorectal cancer (EoCRC) in patients without any family history has been increasing in recent years. Our study leveraged advanced Large Language Models (LLM), like GPT-4, to predict EoCRC in a population comprised of multiple health systems across the United States. There is potential to improve patient care by suggesting early screening for patients who are predicted to be at risk by the model. Methods: We identified a population of 5532 patients aged between 18 and 44 in the Truveta data, which is a collection of 120+ million patient journeys across 30 U.S. health systems. 1376 (24.87%) were diagnosed of CRC based on their ICD and SNOMED-CT codes. Data was split into training (80%) and testing (20%) sets. For the prediction task, we applied GPT-4o and compared with XGBoost, one of the strongest non-generative machine learning models. We used patient demographics (age, gender, race, ethnicity), conditions, and lab results within the last 2 to 7 months prior to CRC diagnosis for model training. The last month before CRC diagnosis was excluded to avoid highly predictive signals. For XGBoost, the demographics were represented as one-hot feature vectors, indicating their presence or absence. Conditions were encoded by their diagnosis frequency within the time frame, while the actual values of the lab results were used as model features. For the LLM model, all patient information was input as plain text. Both conditions and lab results were represented by the names of ICD and LOINC codes in order to capture the clinical context. A Chain-of-Thought prompting strategy, incorporating detailed instructions and CRC-specific knowledge, was employed to guide the LLM. Results: Our test set consisted of 1105 patients in which 279 (25.25%) were diagnosed with CRC. Both XGBoost and GPT-4o achieved comparable results. Despite the uneven distribution of CRC diagnoses in the test samples, the fine-tuned GPT-4o achieved the highest precision (87.43%) and recall (57.35%). This indicates that the model can accurately predict EoCRC in the near future for over half of the patients. Conclusions: This study highlights the potential of using LLM to predict EoCRC in younger population. While the GPT-4o base model contains general medical knowledge, supervised fine-tuning with explicit guideline enhances its predictive capabilities. The high precision of the model performance minimizes the burden of unnecessary screening by identifying patients with relatively high risk. As AI technologies continues to advance, with sufficient governance policy in place, predictive models can be valuable tools for clinicians to suggest early screening and mitigate EoCRC for patients. Model Precision Recall F1 score Accuracy XGBoost 83.77% 57.35% 68.09% 86.43% GPT-4o (base) 68.61% 54.84% 60.69% 82.26% GPT-4o (fine-tuned) 87.43% 57.35% 69.26% 87.15%
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (8)
Wilson Lau
Truveta Inc., Bellevue, WA
Youngwon Kim
School of Biological Sciences, Seoul National University
Sravanthi Parasa
Md Enamul Haque
Sara Daraei
Truveta Inc, Belleue, WA
Rajesh Rao
Truveta Inc, Bellevue, WA
Jay Pillai
Truveta Inc, Bellevue, WA
Anand Oka
Truveta Inc., Bellevue, WA