Evaluating generative AI models for development of patient education quizzes on hormone-related symptoms of cancer therapy.

C Claire Dvorak (Orlando Health, Orlando, FL) J Jeffrey R Smith (Orlando Health, Orlando, FL) N Nikita C. Shah (Orlando Health Cancer Institute, Orlando, FL) P Patrick Kelly (Orlando Health Cancer Institute, Orlando, FL) T Tomas Dvorak

Abstract

e13880 Background: To assess whether generative AI models can effectively design patient-appropriate educational quizzes on hormone-related symptoms of cancer therapy, focusing on quiz structure, question quality, and accuracy. We examined gender-specific content to address distinct educational needs of men and women. Methods: Five current generative AI models (GPT-o1, Claude3.5 Sonnet, Grok2, Gemini1.5, and DeepSeek) were prompted to created two quizzes—one for men and one for women— using information provided to them based on Chapter 3 of the 2024 NCCN Patient Survivorship Guideline: Hormone-Related Symptoms. Each model was to determine number and type of questions (multiple-choice, true/false, and open-ended), categorized into six domains: (1) sex hormones and cancer, (2) type and timing of symptoms, (3) assessment of symptoms, (4) treatment of hot flashes, (5) treatment of gynecomastia, and (6) treatment of urogenital problems. Question quality was assessed using a 5-point Likert scale across four criteria: relevance, clarity, clinical appropriateness, and engagement. Results: Question count per quiz ranged from 5 to 10 (average 7.0), with multiple-choice as the dominant format (63%), followed by true/false (24%) and open-ended (13%). DeepSeek generated the fewest questions (5 each), while Gemini1.5 the most (8 for men, 10 for women). Claude-3.5 received the highest overall score (4.5), followed by GPT-o1 (4.4), DeepSeek (4.2), Gemini1.5 (3.9), and Grok2 (3.8). Across individual criteria, Claude 3.5 ranked highest in relevance (5.0) and clarity (4.8), while GPT-o1 and DeepSeek tied for the highest score in clinical appropriateness (4.8 and 4.6, respectively), and Claude-3.5 (4.0 and GPT-o1 (3.8) scoring best on engagement. All models produced correct answer keys (100% accuracy) based on guideline content, with no "hallucinations" or incorrect answers. We then curated two 10 best-question patient quizzes (female / male) which included model contributions as follows: Claude-3.5 (5/5), GPT-o1 (2/2), Gemini1.5 (1/2), Grok (1/0), DeepSeek (1/1) questions respectively. Conclusions: Generative AI models demonstrated competency in designing structured, clinically relevant, and accurate patient quizzes on hormone-related symptoms during survivorship of cancer care. While the quality of questions varied among models, Claude 3.5 and GPT-o1 consistently delivered the most relevant and engaging content. These findings underscore the potential of AI as a valuable tool in creating patient education materials. Further validation and clinical testing are required to optimize reliability and engagement, as well as patient acceptance and educational impact. A mix of different models may yield the best clinically useful material.

Article Details

Volume / Issue Vol. 43, Issue 16_suppl
Published June 01, 2025
ISSN 0732-183X
Publisher Lippincott Williams & Wilkins

Journal Info

Journal of Clinical Oncology

Lippincott Williams & Wilkins

ISSN: 0732-183X Health Sciences

Authors (5)

C

Claire Dvorak

Orlando Health, Orlando, FL

J

Jeffrey R Smith

Orlando Health, Orlando, FL

N

Nikita C. Shah

Orlando Health Cancer Institute, Orlando, FL

P

Patrick Kelly

Orlando Health Cancer Institute, Orlando, FL

T

Tomas Dvorak