Decoding oncology terminology: Using large language models for patient education.
Abstract
10600 Background: Large language models (LLMs), such as OpenAI's ChatGPT-4, are designed to process natural language and generate responses to text-based prompts. While these models have shown promise in addressing clinical and patient-related inquiries, they lack integration with dedicated medical knowledge databases, leading to potential inaccuracies. Meanwhile, healthcare professionals and models explicitly trained in medical texts frequently rely on specialized terminology, which can create significant barriers to clear and patient-friendly communication. This study aims to utilize LLMs to translate complex medical terminology into easy-to-understand explanations, focusing on hematology and oncology fields where communicating medical concepts to the public is particularly challenging. Our objective is to develop a solution that ensures explanations are both accurate and accessible, bridging the gap between technical medical knowledge and patient comprehension. Methods: We curated a dataset of cancer-related terms and their explanations from two sources: the National Cancer Institute (NCI) Dictionary, which provides detailed medical definitions, and simplified explanations based on National Comprehensive Cancer Network (NCCN) guidelines for patients. Using Meta’s LLaMA 7 B-based chat model, we implemented retrieval-augmented generation (RAG) to enable the model to access the NCI Dictionary as needed. To fine-tune the model to generate patient-friendly explanations, we applied LoRA-based supervised fine-tuning (SFT). The model was evaluated on a holdout set of terms. Readability was measured using the Flesch Reading Ease Score (FRES ) and Dale-Chall Readability Formula (DCRF). Improvements in accessibility were quantified through a two-sample t-test, comparing the mean readability scores of the model's outputs against the baseline. Results: The fine-tuned model demonstrated significant improvements in both accessibility and readability, achieving a 5% and 4% increase in the FRES (baseline: 69.25; output: 72.60, P < 0.01) and DCRF (baseline: 9.50; output: 9.15, P < 0.01), respectively. Preliminary results also indicate the model’s capability to translate entire paragraphs of dense medical text into patient-friendly explanations. Expert validation on a larger scale is currently underway, further solidifying the model's potential to revolutionize patient communication in oncology. This innovative approach sets a new standard for leveraging advanced language models in healthcare education. Conclusions: We successfully trained a LLM specifically designed to simplify complex oncology medical terms into patient-friendly language. By serving as a reliable tool for delivering precise and accessible medical information, it holds the potential to reduce the workload of healthcare providers and enhance patient understanding in clinical settings.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (6)
Yuqing Wang
Inae Park
1Montefiore Medical Center, Internal Medicine, Bronx, United States
Jiahao Peng
State Key Laboratory of Bioinspired Interfacial Materials Science, Institute of Functional Nano & Soft Materials (FUNSOM)
Simo Du
Zhengrui Xiao
1Montefiore Medical Center, Hematology-Oncology, Bronx, United States
Yizhou Chen