Large language model–assisted decision-making in multidisciplinary tumor boards for colorectal cancer.

A Aydan Farzaliyeva (Department of Medical Oncology, Baskent University, Ankara, Turkey) A Arzu Oguz (Department of Medical Oncology, Baskent University, Ankara, Turkey) O Ozden Altundag (Department of Medical Oncology, Baskent University, Ankara, Turkey) Z Zafer Akcali (Department of Medical Oncology, Baskent University, Ankara, Turkey)

Abstract

3521 Background: Evolving colorectal cancer management and rising incidence have increased the workload of multidisciplinary tumor boards (MTBs), placing growing time demands on healthcare professionals. This study evaluates large language models (LLMs) as decision-support tools by comparing AI-generated recommendations with MTB decisions. Methods: This retrospective study used two large language models, Gemini 2.5 (general-purpose) and MedGemma 27B (medical-domain specific), to generate treatment recommendations for patients discussed at MTB. MedGemma 27B was evaluated at two temperature settings (T = 0.0 and T = 1.0). MTB decisions served as the reference standard. Agreement was quantified using Cohen’s kappa, and classification performance was assessed with accuracy, F1 score, and recall. Recommendations were independently reviewed by two MTB members for clinical concordance and safety. Results: A total of 300 colorectal cancer cases were included. Gemini 2.5 demonstrated high agreement with MTB decisions (Cohen’s kappa = 0.792, p < 0.001), with 81.7% full concordance, significantly outperforming MedGemma 27B at both temperature settings (57.0% at T = 0.0 and 62.3% at T = 1.0, p < 0.001 for both). MedGemma 27B showed moderate agreement (k = 0.566 at T = 0.0, k = 0.610 at T = 1.0, p < 0.001), with a modest improvement in concordance at T = 1.0 (p = 0.038). In safety, Gemini 2.5 had the highest rate of clinically safe recommendations (83.7%), outperforming MedGemma at both temperature settings (p < 0.001), with no significant safety difference between MedGemma's temperature settings (p = 0.567). Subgroup analyses showed that Gemini 2.5 had higher discordance in stage IV (p = 0.033), recurrent cases (p < 0.001), and surgical/interventional or active surveillance decisions (p < 0.001). In contrast, MedGemma demonstrated increased discordance in stage II (p < 0.001), ECOG PS ≥2 (p = 0.023–0.036), MSI-unstable tumors (p < 0.001), recurrent disease (p = 0.008–0.010), and surgical/interventional or active surveillance decisions (p < 0.001). Conclusions: Large language models demonstrated meaningful alignment with multidisciplinary tumor board decisions, supporting their potential role in clinical decision support for colorectal cancer, while emphasizing the need for further refinement to ensure consistent and safe clinical integration. Comparison of large language models’ agreement and performance metrics with multidisciplinary tumor board decisions in colorectal cancer. LLM Fully concordant Partially concordant Discordant Accuracy F1 score Recall Cohen's kappa (κ) p value Gemini 2.5 245 (81.7 %) 23 (7.7 %) 32 (10.7 %) 85 % 0.79 0.82 0.792 <0.001** MedGemma 27B temperature 0.0 171 (57.0 %) 64 (21.3 %) 65 (21.7 %) 70.3% 0.64 0.61 0.566 <0.001** MedGemma 27B temperature 1.0 187 (62.3 %) 54 (18.0 %) 59 (19.7 %) 73.5% 0.71 0.66 0.610 <0.001** LLM, large language model.

Article Details

Volume / Issue Vol. 44, Issue 16_suppl
Published June 01, 2026
Pages 3521-3521
ISSN 0732-183X
Publisher Lippincott Williams & Wilkins

Journal Info

Journal of Clinical Oncology

Lippincott Williams & Wilkins

ISSN: 0732-183X Health Sciences

Authors (4)

A

Aydan Farzaliyeva

Department of Medical Oncology, Baskent University, Ankara, Turkey

A

Arzu Oguz

Department of Medical Oncology, Baskent University, Ankara, Turkey

O

Ozden Altundag

Department of Medical Oncology, Baskent University, Ankara, Turkey

Z

Zafer Akcali

Department of Medical Oncology, Baskent University, Ankara, Turkey