Image harmonization for PD-(L)1 immune checkpoint inhibitor response prediction using radiomic features and deep learning in advanced NSCLC.
Abstract
e13689 Background: Machine learning models using radiomic features have demonstrated predictive and prognostic value in oncology applications. One challenge in translating these models into clinical use is the wide variation in imaging protocols across clinical sites. Variations in reconstruction kernels in CT images can alter radiomic features, introducing bias and impacting model generalizability. We investigated the effects of a physics-based post-processing harmonization method on (1) pre-defined radiomic features, and (2) the performance of a deep learning model to predict immune checkpoint inhibitor (ICI) response in advanced NSCLC. Methods: This study was performed on an internally curated, multi-institutional real-world dataset of chest CT scans from 1,188 advanced NSCLC patients treated with PD-(L)1 ICIs in academic and community settings from the US and Europe. The dataset included images spanning eight manufacturers, over 100 scanner models, and 31 reconstruction kernels. A post-processing harmonization algorithm, leveraging phantom-derived protocol libraries and an online noise power spectrum estimation method, was used to standardize images to a common kernel, correcting for reconstruction variability. In the first investigation, 85 radiomic features (e.g., shape, texture) were extracted from the native images and from images harmonized to a common kernel and compared. Performance of the deep learning model in predicting ICI therapy response was quantified by the area under the ROC curve for Progression-Free Survival at 6 months (PFS6 AUC) and Overall Survival at 12 months (OS12 AUC), comparing models trained with and without harmonization. The models were evaluated on an independent test set of 92 patients who received ICI in the first line as a monotherapy from an institution not used for training. Results: Variability in radiomic features due to convolution kernels was significant; the mean difference between images reconstructed by the Lung kernel compared to the Standard kernel was 60%, with a maximum of 359%. After Lung kernel images were harmonized to match the appearance of the Standard images, the mean difference was reduced to 6%, with a maximum of 23%. The deep learning model performance improved when trained on harmonized images, increasing PFS6 AUC from 0.64 to 0.70 and OS12 AUC from 0.75 to 0.79. Conclusions: The large variation in reconstruction kernels used in clinical CT imaging practice may confound radiomic features and compromise the performance of predictive models. Post-processing image harmonization successfully reduced these variations, reducing the bias in radiomic features and improving the predictive performance of a deep learning model for ICI therapy response. This study demonstrates the importance of harmonization techniques to ensure imaging biomarkers that are generalizable.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (4)
Taly Schmidt
Onc.AI, San Carlos, CA
Roshan J Vincent
Onc.AI, San Carlos, CA
Chiharu Sako
Petr Jordan
Onc.AI, San Carlos, CA