Prompt engineering and readability of large language model–generated patient education materials in hematology and oncology.
Abstract
e13668 Background: Patients with cancer must understand complex diagnostic and treatment concepts to make informed healthcare decisions, which disadvantages individuals with limited health literacy. To address this, the Agency for Healthcare Research and Quality and the National Institutes of Health recommend that patient-facing materials be written at a fifth- to sixth-grade reading level. However, most oncology education materials exceed this threshold. Large language models (LLMs) are increasingly used to generate patient-facing educational content and, with appropriate prompt engineering, may be able to produce materials at recommended readability levels across hematologic and solid malignancies. Methods: We generated patient-facing educational explanations for five malignancies: acute myeloid leukemia, diffuse large B-cell lymphoma, breast cancer, colorectal cancer, and metastatic lung cancer. Three LLMs were queried using predefined prompt engineering techniques with varying instructional constraint, including general prompts, explicit grade-level–targeted prompts, meta-generated prompts, and teach-back prompts. Readability of English-language outputs was assessed using the Simple Measure of Gobbledygook (SMOG) and Flesch–Kincaid Grade Level (FKGL) via a Python-based text analysis pipeline. The primary analytic goal was to evaluate whether prompt structure shifted outputs toward or below the recommended sixth-grade reading level. Differences across prompt types were evaluated using one-way ANOVA with post-hoc Tukey HSD testing. Results: Prompt type significantly influenced readability for both SMOG (F = 121.2, p < 0.001) and FKGL (F = 81.2, p < 0.001). General prompts produced the least readable outputs. Grade-level–targeted, meta-generated, and teach-back prompts significantly improved readability compared with general prompts (p < 0.001). There were no significant differences in SMOG or FKGL between grade-level–targeted, meta, and teach-back prompts. Persona-style prompts produced higher reading grade levels than constrained and meta prompts. No prompt consistently achieved a sixth-grade reading level by SMOG, though some models reached this threshold by FKGL. Conclusions: Prompt structure meaningfully affects the readability of LLM-generated oncology patient education materials. Explicitly constrained and meta-generated prompts improve readability, whereas persona-based prompting may increase textual complexity. These findings offer practical guidance for clinicians using LLMs to support accessible, patient-centered cancer education.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (3)
Naseeruddin Naseem
Tulane University School of Medicine, New Orleans, LA
Husayn Ramji
Tulane University School of Medicine, New Orleans, LA
Jeffrey Wiese
Tulane University Schoo of Medicine, New Orleans, LA