Towards personalized management of intraductal papillary mucinous neoplasms with multimodal artificial intelligence.
Abstract
e16461 Background: According to the Fukuoka and Kyoto consensus guidelines, the clinical management of pancreatic intraductal papillary mucinous neoplasms (IPMNs) primarily depends on imaging features, cytology, and clinical variables such as CA19-9 serum levels. While these guidelines are sensitive to high-grade or invasive IPMN, they lack specificity, resulting in surgical overtreatment. We propose a clinically applicable multimodal deep learning model to accurately predict the optimal management approach, i.e., surgical resection or surveillance, based on imaging and clinical data. Methods: This retrospective study included 180 patients with IPMN who underwent surgical resection at Indiana University School of Medicine. We developed prediction models for the most optimal management - surgical resection (for high-grade and invasive IPMN) vs surveillance (for low-grade IPMN) - based on individual preoperative magnetic resonance imaging (MRI) sequences (T1 with and without contrast, T2) and clinical data (age, sex, race, BMI, CA19-9, cyst size, duct dilation, IPMN subtype, family history, pancreatitis, unintentional weight loss). In addition, we developed multimodal models integrating MRI sequences in an early fusion manner, using ResNet-34 architecture. Pancreatic region of interest (ROI) was segmented using nnUNet, trained for pancreas segmentation in multiple MRI sequences. Model performance was evaluated through 5-fold cross-validation, and model performance on an independent holdout test set was compared with the clinical gold standard performance based on the Fukuoka consensus guidelines using F1-score and the area under the receiver-operating characteristic curve (AUC). Results: The multimodal model trained on both imaging and clinical data (F1-score: 0.83, 95% CI: 0.61, 1.00; AUC: 0.93, 95% CI: 0.60, 1.00) outperformed the multimodal model trained on imaging data only (F1-score: 0.71, 95% CI: 0.45, 0.91; AUC: 0.87, 95% CI: 0.70, 1.00) as well as unimodal models trained on T2-weighted MRIs alone (F1-score: 0.59, 95% CI: 0.42, 0.81; AUC: 0.70, 95% CI: 0.44, 0.94), T1-weighted MRIs alone (F1-score: 0.58, 95% CI: 0.40, 0.76; AUC: 0.67, 95% CI: 0.39, 0.92), and T1-weighted contrast-enhanced MRIs alone (F1-score: 0.55, 95% CI: 0.37, 0.74; AUC: 0.47, 95% CI: 0.14, 0.80) on the hold-out test set. In addition, the multimodal model outperformed the stratification performance of the Fukuoka criteria (F1-score: 0.54, 95% CI: 0.34, 0.73; AUC: 0.58, 95% CI: 0.31, 0.82) on the holdout test set, showcasing the clinical potential of multimodal AI to personalize management in patients with IPMN. Conclusions: Our findings highlight the added value of multimodal integration, suggesting that multimodal AI models can serve as a valuable diagnostic tool for individualized clinical management of IPMN. External validation is planned to bridge the transition into clinical practice.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (7)
Muhammad Ibtsaam Qadir
Weldon School of Biomedical Engineering, Purdue University, West Lafayette, IN
Jackson Baril
Department of Surgical Oncology, Indiana University School of Medicine, Indianapolis, IN
Duane Schonlau
Department of Radiology, Indiana University School of Medicine, Indianapolis, IN
Michele T. Yip-Schneider
Department of Surgical Oncology, Indiana University School of Medicine, Indianapolis, IN
Thi Thanh Thoa Tran
Department of Surgical Oncology, Indiana University School of Medicine, Indianapolis, IN
Christian Schmidt
Medical Research Council Prion Unit at University College London, University College London Institute of Prion Diseases
Fiona R. Kolbinger