Artificial intelligence‑driven virtual tumor board enhances precision care in myelodysplastic syndromes (MDS)

D David Swoboda (1Tampa General Hospital Cancer Institute, Tampa, United States) A Amy DeZern (1Johns Hopkins University School of Medicine, Oncology, Baltimore, United States) J James England (3Sunnybrook Health Sciences Centre, Toronto, Canada) S Sangeetha Venugopal (4University of Miami, Miami, United States) T Thomas Kehoe (5University of South Florida College of Medicine, Tampa, United States) B Brandon Aubrey (6Massachusetts General Hospital, Boston, United States) M Marco Gabriele Raddi (5University of Florence, MDS Unit, Florence, Italy) A Angela Consagra (30University of Florence, MDS Unit, Hematology, AOU Careggi - Department of Experimental and Clinical Medicine, Florence, Italy) J Jiasheng Wang (Dr. Li Dak Sum and Yip Yio Chin Center for Stem Cells and Regenerative Medicine, Zhejiang University School of Medicine) G Gustavo Rivero (3Tampa General Hospital Cancer Institute, Tampa, United States) M Maximilian Stahl A Amer Zeidan (18Yale School of Medicine - Yale Cancer Center, New Haven, United States) T Torsten Haferlach (7Munich Leukemia Laboratory, Munich, Germany) A Andrew Brunner (3Dana-Farber Cancer Institute, Boston, United States) R Rena Buckstein V Valeria Santini (7DMSC University of Florence, AOUC, MDS Unit, Hematology, Florence, Italy) M Matteo Della Porta (1IRCCS Humanitas Research Hospital, AI Center, Rozzano, Italy) M Mikkael Sekeres (13Sylvester Cancer Center, University of Miami Health System, Miami, United States) A Aziz Nazha (1Department of Medical Oncology, Sidney Kimmel Cancer Center, Thomas Jefferson University, Philadelphia, PA)

Abstract

Abstract Background and Aims Myelodysplastic syndromes (MDS) present significant diagnostic and therapeutic challenges, often requiring input from multidisciplinary teams and subspecialty-trained leukemia and pathology experts. The complexity is compounded by evolving classification systems (e.g., WHO and ICC) and an expanding therapeutic landscape, which complicate real-time clinical decision-making. Although large language models (LLMs) such as ChatGPT have demonstrated promise in medical domains, they frequently yield inaccurate or overly generalized responses when applied to complex hematologic scenarios. Even state-of-the-art models—capable of human-like reasoning and sometimes outperforming clinicians in general tasks—have not been systematically evaluated in the context of advanced hematologic disorders such as MDS. To address this gap, we first assessed the performance of leading LLMs, including ChatGPT, Claude, and DeepSeek, on challenging, real-world MDS cases. After identifying their limitations, we built the Virtual MDS Panel (VMP), a coordinated AI system in which AI agents—task-bound software assistants that understand natural language and help users complete tasks, answer questions, and make decisions efficiently—are trained on domain knowledge (WHO/ICC; IPSS-R/IPSS-M; NCCN) and explicit decision rules to collaborate and produce tumor-board–level recommendations. Methods VMP comprises four specialized AI agents: a moderator agent that receives clinical queries; a pathology agent trained on WHO/ICC criteria; a prognostication agent using risk models (IPSS, IPSS-R, IPSS-M) and a therapy agent grounded in NCCN and ELN guidelines. For each case, a physician submits a clinical scenario to the moderator, which breaks down the query and delegates sub-tasks to relevant agents. The moderator then synthesizes their responses into a structured output. To evaluate performance, we created a test set of 30 complex, real-world MDS cases. VMP responses were compared to those from leading LLMs (ChatGPT-4o, GPT-o3, Claude, DeepSeek). Eleven international MDS experts, blinded to response sources, independently scored outputs for accuracy, clinical relevance, and completeness (Likert scale 1–5). They also assessed diagnostic reasoning, prognostic validity, and treatment recommendations, while classifying factual errors as none, minor, or major. To evaluate the consistency of expert ratings, we used the intraclass correlation coefficient (ICC) to measure how well experts agreed on numerical scores and Cohen's κ (kappa) to assess their agreement when identifying errors. Results VMP achieved an overall expert-rated accuracy of 93%, outperforming GPT-o3 (82%), GPT-4o (80%), DeepSeek (71%), and Claude (66%). When examining each domain, the VMP overall scored highest 4.2 (mean on scale from 1-5 among all experts): Diagnosis 4.3, Prognosis 4.4 and therapy selection 3.9, compared to GPT-o3, at 3.6 overall, 3.7 / 3.6 / 3.6, GPT-4o (3.2 overall) 3.1 / 3.2 / 3.4, DeepSeek (3.0) 2.9 / 3.0 / 3.1, and Claude (2.9) 2.7 / 2.9 / 3.1, respectively. According to expert opinion, the VMP registered just 9% major factual errors compared to GPT-o3 26%, GPT-4o 26%, DeepSeek 33%, and Claude 36%. Minor factual errors were present at 36% in VMP vs 47-52% in the other 4 models. Experts showed strong agreement in their evaluations, with high consistency in scoring (ICC = 0.81) and in identifying AI errors or hallucinations (κ = 0.76), confirming the reliability of the review process.Conclusions We developed an advanced, MDS-focused AI system that increased accuracy and alignment with expert practice, outperforming current state-of-the-art general AI models. By emulating a virtual tumor board, the system offers structured, evidence-based guidance that can aid hematologists in optimizing their precision care.

Article Details

Journal Blood
Volume / Issue Vol. 146, Issue Supplement 1
Published November 03, 2025
Pages 7363-7363
ISSN 0006-4971
Publisher Elsevier BV

Journal Info

Blood

Elsevier BV

ISSN: 0006-4971 Health Sciences

Authors (19)

D

David Swoboda

1Tampa General Hospital Cancer Institute, Tampa, United States

A

Amy DeZern

1Johns Hopkins University School of Medicine, Oncology, Baltimore, United States

J

James England

3Sunnybrook Health Sciences Centre, Toronto, Canada

S

Sangeetha Venugopal

4University of Miami, Miami, United States

T

Thomas Kehoe

5University of South Florida College of Medicine, Tampa, United States

B

Brandon Aubrey

6Massachusetts General Hospital, Boston, United States

M

Marco Gabriele Raddi

5University of Florence, MDS Unit, Florence, Italy

A

Angela Consagra

30University of Florence, MDS Unit, Hematology, AOU Careggi - Department of Experimental and Clinical Medicine, Florence, Italy

J

Jiasheng Wang

Dr. Li Dak Sum and Yip Yio Chin Center for Stem Cells and Regenerative Medicine, Zhejiang University School of Medicine

G

Gustavo Rivero

3Tampa General Hospital Cancer Institute, Tampa, United States

M

Maximilian Stahl

A

Amer Zeidan

18Yale School of Medicine - Yale Cancer Center, New Haven, United States

T

Torsten Haferlach

7Munich Leukemia Laboratory, Munich, Germany

A

Andrew Brunner

3Dana-Farber Cancer Institute, Boston, United States

R

Rena Buckstein

V

Valeria Santini

7DMSC University of Florence, AOUC, MDS Unit, Hematology, Florence, Italy

M

Matteo Della Porta

1IRCCS Humanitas Research Hospital, AI Center, Rozzano, Italy

M

Mikkael Sekeres

13Sylvester Cancer Center, University of Miami Health System, Miami, United States

A

Aziz Nazha

1Department of Medical Oncology, Sidney Kimmel Cancer Center, Thomas Jefferson University, Philadelphia, PA