Diagnostic accuracy of machine learning models in glioma classification: A meta-analysis.
Abstract
2079 Background: Machine learning (ML) is promising in IDH-based glioma classification using magnetic resonance imaging (MRI), but variability in methods and algorithms necessitates a comprehensive evaluation. This meta-analysis assesses the pooled diagnostic performance of ML-based approaches. Methods: A literature search was conducted in January 2025 across PubMed, MEDLINE, and Cochrane. MICCAI, RSNA, and SNO meeting abstracts were additionally reviewed. Eligible studies evaluating ML models for IDH-based glioma classification using MRI were included. Data were pooled using a random-effects model, analyzing sensitivity, specificity, heterogeneity, and publication bias via Egger’s test and funnel plots. Leave-one-out analysis was conducted. Results: A total of 5982 cases were analyzed. Gliomas were classified as WHO Grades II (11.5%), III (23.1%), and IV (65.4%). Histopathology-based reference standards, including genetic and molecular testing, were used in 73.1% of studies, while immunohistochemistry, pathology, biopsy-proven markers, and immunohistopathologic diagnosis were each used in 3.8–7.7%. Deep learning models, including CNNs and ResNet, were the most used classifiers (30.8%), followed by Support Vector Machines (26.9%). Ensemble methods, such as Random Forest accounted for 19.2%, regression-based approaches (LASSO, logistic regression) for 15.3%, and other techniques like multilayer perceptron and AdaBoost for 7.7%. This meta-analysis included 25 studies for sensitivity and 26 for specificity, using a random-effects model with DerSimonian-Laird estimation. Pooled sensitivity was 83.0% (95% CI: 79.5–86.5%) and specificity was 78.6% (95% CI: 73.7-83.4%), both statistically significant (p < 0.0001). Substantial heterogeneity was found (I² = 100% for both), with Cochran’s Q values of 1.9e+06 for sensitivity and 4.2e+06 for specificity (p < 0.001). Leave-one-out analysis showed minimal variation in pooled estimates (sensitivity: 82.6–83.8%, specificity: 77.6–79.4%). Egger’s test revealed significant small-study effects (p = 0.0001 for both), suggesting potential publication bias. Conclusions: ML models demonstrated moderate diagnostic performance in IDH-based glioma classification, achieving a sensitivity of 83.1% and specificity of 78.6%. However, substantial heterogeneity and potential biases pose significant challenges to their clinical implementation. To enhance the reliability and broader applicability of ML models in IDH-based glioma diagnosis, standardization of imaging protocols and external validation are imperative. Meta-analytical findings. Metric Sensitivity (%) Specificity (%) Pooled Estimate 83.0 (79.5–86.6) 78.6 (73.7–83.4) Z-Value 46.4 31.76 P-Value <0.0001 <0.0001 I² (%) 100 100 T² 80.0 159.0 Cochran's Q 1.9e+06 (df=24, p<0.001) 4.2e+06 (df=25, p<0.001) Leave-One-Out Range 82.6–83.8 77.6–79.4 Egger’s Test (P-value) 0.001 0.001
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (11)
Maya Gowda
Cornell University, New York, NY
Zouina Sarfraz
Khalis Mustafayev
Miami Cancer Institute, Baptist Health South Florida, Miami, FL
Logan Spencer Spiegelman
Miami Cancer Institute, Baptist Health South Florida, Miami, FL
Fatma Nihan Akkoc Mustafayev
Miami Cancer Institute, Baptist Health South Florida, Miami, FL
Mohammad Arfat Ganiyani
6Miami Cancer Institute, Miami, United States
Michael W. McDermott
Arun Maharaj
Rupesh Kotecha
Yazmin Odia
Manmeet Singh Ahluwalia
Miami Cancer Institute, Baptist Health South Florida, Miami, FL