Concordance between a multidisciplinary tumor board and ChatGPT 4.0 in the management of genitourinary tumors: A real-world analysis.
Abstract
e23256 Background: Artificial intelligence (AI) is emerging as a valuable tool to support clinical decisions in oncology. Nevertheless, its practical implementation in clinical settings requires further validation. This study evaluates the concordance between a Genitourinary Multidisciplinary Tumor Board (MTB) and ChatGPT 4.0. Methods: An observational, retrospective design was used to analyze cases discussed in the MTB at a university hospital in Mexico from January 2023 to January 2024. The protocol was developed before implementation, and formal approval was received from the Ethics Committee. Cases with a complete medical record and a comprehensive tumor board record were included. MTB recommendations were compared with those from ChatGPT 4.0 using an open-ended approach. ChatGPT received a structured clinical prompt (demographics, history and physical, laboratory results, imaging, and pathology) requesting a management plan. Concordance between plans was analyzed across three levels of decision-making: Level I (Diagnostic vs. Therapeutic Approach), Level II (Type of Management), specifying tests or treatments, and Level III (Comprehensive Management), focusing on specific tests or treatment selection. Concordance rates and Cohen’s kappa coefficients were calculated and interpreted using medical standards: κ > 0.80 = excellent, κ 0.61–0.80 = acceptable, and κ < 0.60 = questionable. Results: Thirty-seven patients with genitourinary tumors met inclusion criteria with complete medical and MTB records. Median age was 64, 78% (n = 29) were male and 22% female (n = 8). Tumor types included kidney (27%, n = 10), testicular (27%, n = 10), bladder (19%, n = 7), penile (16%, n = 6), prostate (3%, n = 1), and adrenal (3%, n = 1). Concordance between the multidisciplinary tumor board and ChatGPT 4.0 was assessed across three decision levels. Level I showed 86% concordance (κ = 0.65, acceptable). Level II had 70% concordance (κ = 0.61, acceptable), while Level III dropped to 54% (κ = 0.59, questionable). Conclusions: Managing clinical oncology cases is complex, often requiring input from multidisciplinary teams. During MTB discussions, diagnostic results are reviewed, and team feedback may lead to adjustments in treatment plans. This real-time reevaluation remains beyond the scope of large language models like ChatGPT. Exploratory analysis found that most discordance occurred in high-complexity cases requiring dynamic data integration. Further study is essential to identify potential factors influencing differences between MTB and ChatGPT recommendations. This understanding may enhance its role as a supportive tool in oncology. Concordance rates and kappa values across clinical decision-making levels. Level Concordance (%) Kappa Interpretation Level I 86.49 0.65 Acceptable Level II 70.27 0.61 Acceptable Level III 54.05 0.59 Questionable
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (6)
Magda Arredondo
Hospital Universitario "Dr. José Eleuterio González", Monterrey, NL, Mexico
Omar A. Zayas
Hospital Universitario "Dr. José Eleuterio González", Monterrey, Mexico
Christopher Cerda-Contreras
Hospital Universitario "Dr. José Eleuterio González", Monterrey, Mexico
Fernando J. Peña-González
Hospital Universitario "Dr. José Eleuterio González", Monterrey, Mexico
Carlos Eduardo Salazar-Mejía
Hospital Universitario "Dr. José Eleuterio González", Monterrey, Mexico
Oscar Vidal-Gutierrez
Hospital Universitario Dr. Jose Eleuterio Gonzalez, Monterrey, NL, Mexico