Performance of ChatGPT in generating patient-facing cancer survivorship care plans.
Abstract
1673 Background: Survivorship care plans (SCPs) summarize cancer treatments and provide evidence-based recommendations for surveillance, screening, and management of treatment-related complications. Despite the availability of guidelines and templates, SCPs remain underutilized in routine practice due to the time-intensive and error-prone nature of manual creation. Empowering cancer survivors to use publicly available large language models (LLM)–based chatbots, such as ChatGPT, may offer a feasible solution to this problem. Methods: Standardized patient profiles were developed for six common cancers: breast, colorectal, non–small cell lung, small cell lung, prostate cancer, and diffuse large B-cell lymphoma. Profiles were reviewed by a primary care physician and an oncologist to ensure clinical accuracy and representativeness of common survivorship scenarios. Using lay language to simulate real-world patient input, ChatGPT (version 5.2, publicly available at the time of the study) was prompted with each profile using the question, “What follow-up care do I need?” Common survivorship symptoms for each cancer were also tested. Responses were evaluated across five domains: medical accuracy (concordance with current ASCO or NCCN guidelines), thoroughness, patient safety and risk framing, clarity and patient comprehension, and actionability. Five-point Likert scales were used (1 = not at all; 5 = very much). Scores were averaged across two independent physician raters. Readability was assessed using the Flesch Reading Ease Score. Results: ChatGPT-generated SCPs demonstrated high medical accuracy (mean score 4.25/5), with no factual errors or hallucinations identified on manual review. Patient safety and risk framing scored 3.75/5; symptom red flags and care escalation guidance were appropriately stated in most scenarios. Thoroughness was moderate (3.50/5), with key survivorship elements, such as smoking cessation counseling or genetic risk considerations; sometimes omitted unless explicitly prompted. Readability was limited, with a mean Flesch Reading Ease Score of 40.5, corresponding to a college reading level and exceeding recommended readability standards for patient education materials. Clarity (3.38/5) and actionability (3.38/5) were similarly constrained due to dense medical language. Conclusions: ChatGPT demonstrated high medical accuracy in generating survivorship care plans across six common cancers, supporting the feasibility of using publicly available LLMs to guide survivorship care. However, improvements in comprehensiveness, readability, and patient-centered actionability are necessary before such tools can be safely integrated into clinical survivorship workflows.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (4)
Kaili Du
Cayuga Medical Center at Ithaca, Ithaca, NY
Yuting Zhang
Shenzhen Crystalo Biopharmaceutical Co., Ltd., Shenzhen, Guangdong, China.
Kristina Pradhan
Cayuga Medical Center, Ithaca, NY
Chenyu Sun