Large language model–assisted decision-making in multidisciplinary tumor boards for colorectal cancer.
Abstract
3521 Background: Evolving colorectal cancer management and rising incidence have increased the workload of multidisciplinary tumor boards (MTBs), placing growing time demands on healthcare professionals. This study evaluates large language models (LLMs) as decision-support tools by comparing AI-generated recommendations with MTB decisions. Methods: This retrospective study used two large language models, Gemini 2.5 (general-purpose) and MedGemma 27B (medical-domain specific), to generate treatment recommendations for patients discussed at MTB. MedGemma 27B was evaluated at two temperature settings (T = 0.0 and T = 1.0). MTB decisions served as the reference standard. Agreement was quantified using Cohen’s kappa, and classification performance was assessed with accuracy, F1 score, and recall. Recommendations were independently reviewed by two MTB members for clinical concordance and safety. Results: A total of 300 colorectal cancer cases were included. Gemini 2.5 demonstrated high agreement with MTB decisions (Cohen’s kappa = 0.792, p < 0.001), with 81.7% full concordance, significantly outperforming MedGemma 27B at both temperature settings (57.0% at T = 0.0 and 62.3% at T = 1.0, p < 0.001 for both). MedGemma 27B showed moderate agreement (k = 0.566 at T = 0.0, k = 0.610 at T = 1.0, p < 0.001), with a modest improvement in concordance at T = 1.0 (p = 0.038). In safety, Gemini 2.5 had the highest rate of clinically safe recommendations (83.7%), outperforming MedGemma at both temperature settings (p < 0.001), with no significant safety difference between MedGemma's temperature settings (p = 0.567). Subgroup analyses showed that Gemini 2.5 had higher discordance in stage IV (p = 0.033), recurrent cases (p < 0.001), and surgical/interventional or active surveillance decisions (p < 0.001). In contrast, MedGemma demonstrated increased discordance in stage II (p < 0.001), ECOG PS ≥2 (p = 0.023–0.036), MSI-unstable tumors (p < 0.001), recurrent disease (p = 0.008–0.010), and surgical/interventional or active surveillance decisions (p < 0.001). Conclusions: Large language models demonstrated meaningful alignment with multidisciplinary tumor board decisions, supporting their potential role in clinical decision support for colorectal cancer, while emphasizing the need for further refinement to ensure consistent and safe clinical integration. Comparison of large language models’ agreement and performance metrics with multidisciplinary tumor board decisions in colorectal cancer. LLM Fully concordant Partially concordant Discordant Accuracy F1 score Recall Cohen's kappa (κ) p value Gemini 2.5 245 (81.7 %) 23 (7.7 %) 32 (10.7 %) 85 % 0.79 0.82 0.792 <0.001** MedGemma 27B temperature 0.0 171 (57.0 %) 64 (21.3 %) 65 (21.7 %) 70.3% 0.64 0.61 0.566 <0.001** MedGemma 27B temperature 1.0 187 (62.3 %) 54 (18.0 %) 59 (19.7 %) 73.5% 0.71 0.66 0.610 <0.001** LLM, large language model.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (4)
Aydan Farzaliyeva
Department of Medical Oncology, Baskent University, Ankara, Turkey
Arzu Oguz
Department of Medical Oncology, Baskent University, Ankara, Turkey
Ozden Altundag
Department of Medical Oncology, Baskent University, Ankara, Turkey
Zafer Akcali
Department of Medical Oncology, Baskent University, Ankara, Turkey