The efficacy and limitations of artificial intelligence (AI) in radiologic tumor assessment.
Abstract
e20002 Background: Numerous Artificial Intelligence (AI) models for evaluating radiologic images are emerging, yet limited data exist to demonstrate the efficacy of these tools, particularly within oncology. Notably, these tools often require assistance from radiologists to verify their accuracy, and thus further research is needed to help characterize the current efficacy and limitations of AI in radiologic tumor assessment. This study evaluated ChatGPT-5.2 in thoracic tumor identification and measurement in order to enhance our understanding of the current efficacy and limitations of AI in oncologic radiology assessments. Methods: 61 publicly available CT scans with known thoracic tumors from the Lung CT Diagnosis collection (The Cancer Imaging Archive) were analyzed, with 1 dominant tumor per scan. A single radiologist identified tumors using a representative slice and measured maximal anteroposterior (AP), transverse (TV), and craniocaudal (CC) diameters, with tumor volume calculated using the ellipsoid formula. We sequentially prompted ChatGPT-5.2 to identify tumors on axial and coronal images and to generate corresponding measurements using DICOM metadata. Radiologist and AI measurements were compared using paired t-tests. Results: ChatGPT showed no significant differences from radiologist measurements for axial anteroposterior (26.2 vs 25.2 mm, p=0.176) or transverse diameters (26.9 vs 26.2 mm, p=0.397) across 61 tumor assessments. In contrast, craniocaudal diameters (34.6 vs 24.4 mm, p<0.001) and derived tumor volumes (21.2 vs 14.4 mL, p<0.001) were significantly overestimated. Average percent differences were modest for AP (7.3%) and TV (5.0%) but markedly higher for CC (51.5%) and volume (79.0%). Percent differences for volume measurements ranged from -100% to 745.6% different from the radiologist assessment. The median number of prompts required for correct tumor identification by ChatGPT was 3 for both axial and coronal images. Conclusions: Although ChatGPT-5.2 is not a radiology-specific AI model, this tool utilizes similar approaches to radiology-specific models and can be trained to assess images. While the average percentage differences between ChatGPT and radiologist-assessed images appeared similar, we found wide variation in assessment percentage from image to image. Findings highlight that although AI models may provide a mechanism for aggregating and analyzing large volumes of data, further work is needed to identify how best to incorporate AI tools in radiologic tumor assessment. Comparison of radiologist and ChatGPT-5.2 tumor measurements (n = 61). Measurement Radiologist Mean ChatGPT Mean p -value (paired t -test) Mean % Difference % Difference Range AP (mm) 25.2 26.2 0.176 7.3% -38.9 to 75% TV (mm) 26.2 26.9 0.397 5.0% -65.4 to 120% CC (mm) 24.4 34.6 <0.001 51.5% -100 to 380% Tumor Volume (mL) 14.4 21.2 <0.001 79.0% -100 to 745.6%
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (5)
Sean Michael King
University of Oklahoma College of Medicine, Oklahoma City, OK
Nirmal Choradia
Stephenson Cancer Center, The University of Oklahoma Health Sciences Center, Oklahoma City, OK
Ryan David Nipp
Stephenson Cancer Center, The University of Oklahoma Health Sciences Center, Oklahoma City, OK
Austin Crose
University of Oklahoma College of Medicine, Oklahoma City, OK
Daniel Zhao
University of Oklahoma Hudson College of Public Health, Oklahoma City, OK