Assessing the performance of ChatGPT in addressing ethical dilemmas in oncology.
Abstract
e23292 Background: Large language models such as ChatGPT have been shown to demonstrate accuracy in diagnostic reasoning and treatment plan generation.1,2 However, its utility in navigating ethically complex medical scenarios is less well explored.3 Herein, we evaluated the performance of GPT-5 in addressing commonly encountered ethical cases in oncology. Methods: We conducted a descriptive, cross-sectional approach to assess GPT-5 performance in response to eight ethically complex medical vignettes. A panel of eight board-certified oncologists independently assessed the LLM-generated responses. Vignettes and evaluation dimensions were reviewed by a professor of bioethics who serves as both a clinical ethicist and an institutional board (IRB) member to establish content validity prior to clinician assessment. Outcomes included ethical relevance, reasoning, accuracy, practicability, comprehensibility, and completeness. Oncologists rated each dimension on a 5-point Likert-type scale (1 = strongly disagree to 5 = strongly agree). Interrater reliability was assessed using intraclass correlation coefficient. Results: Of the six outcomes used to evaluate GPT-5 response, the mean score was 4.31/5. GPT-5 responses were rated highest in comprehensibility, defined as the response being clearly written and easy to understand (4.55/5). It was rated lowest in completeness, defined as the response fully addressing all ethical considerations of the case (3.97/5). An intraclass coefficient (ICC) of 0.83 indicates good interrater reliability. GPT-5 response for each clinical vignette was also assessed. Responses were rated highest for the ethical dilemma concerning informed consent, which assessed whether a patient with previously alert, articulated wishes who had developed recent altered mental status had capacity to consent for GBM chemotherapy (4.58/5). Responses were rated lowest for the ethical dilemma of non-disclosure of diagnosis, which assessed approaches to a family hesitant about a clinician sharing a cancer diagnosis with a patient (3.98/5). No significant difference in GPT-5 performance across dimensions or clinical vignettes was found. Conclusions: GPT-5 can provide comprehensible, mostly accurate, and mostly practical responses to help clinicians navigate common ethical dilemmas, although it may not fully capture all ethical considerations. Further research is needed to understand how LLMs can be effectively applied to support ethical decision-making in clinical practice.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (3)
Christabelle Junaidi
UCSF, San Francisco, CA
Anita Ho
UCSF, San Francisco, CA
Brian Schulte
UCSF Helen Diller Family Comprehensive Cancer Center, San Francisco, CA