Evaluation of fine-tuned small language models (SLMs) vs large language models (LLMs) in modeled patient conversations for lung cancer screening.

S Sanjay Khanna (The Royal Marsden Hospital, London, United Kingdom) H Hitesh Khanna (National Health Service, London, United Kingdom) Y Yueqi Ge (The Royal Marsden Hospital, London, United Kingdom) G Gina Sherpa (Imperial College Healthcare NHS Trust, London, United Kingdom) R Richard Lee

Abstract

e24116 Background: Lung Cancer Screening (LCS) uptake in the US remains suboptimal at approximately 18%, with disproportionately lower rates among those with limited health literacy. While Generative AI offers a scalable solution to improve patient access to information, uninstructed commercial large language models (LLMs) exhibit language complexity that limits their utility as equitable decision-support tools. Building on our prior work demonstrating their limitations, we aimed to develop a deployable, patient-centric, small language model (SLM) to address this accessibility gap. Methods: We fine-tuned two SLMs (Google: Gemma-7B and MedGemma-4B) using Parameter-Efficient Fine-Tuning (PEFT/LoRA) on a curated dataset of physician-scored, empathetic dialogues. We modelled conversations using five patient personas (e.g. anxious, low health-literacy) to compare fine-tuned SLMs against uninstructed LLMs (GPT-4o, Gemini Pro 2.5), with static patient information (FAQs). Endpoints included readability (Flesch-Kincaid Grade Level [FKGL]), empathy (reduction in "fear-inducing" sentiment via RoBERTa analysis), clinical safety (semantic adherence to national patient information guidance via BERTScore) and technical density (Type-Token Ratio). Results: Uninstructed LLMs produced responses with a mean reading difficulty of Grade 11.8 (SD 1.7), exceeding the recommended 6th-8th grade level for patient materials. In contrast, the fine-tuned Gemma-7B SLM achieved a mean Grade Level of 6.6 (SD 2.2), making it significantly more accessible than both the generalist models ( P < .001) and static patient information ( P < .001). Regarding empathy, the 7B model reduced the frequency of fear-inducing language by 44% compared to the LLMs (12.1% vs 21.6%; P < .05). Crucially, this simplification did not compromise safety; the 7B model demonstrated high semantic adherence to official patient guidance (BERTScore 0.95) and was non-inferior to public FAQs ( P > 0.05). The smaller MedGemma-4B, however, exhibited a trade-off: while highly accurate, it prioritised technical density (Type-Token Ratio 0.97) over readability (Grade 14.8; P < .001 vs 7B), rendering it less suitable for direct patient interaction. Conclusions: Building on our prior demonstration of an equity blindspot with uninstructed commercial LLMs, this study suggests that a fine-tuned 7B-parameter SLM may decouple clinical accuracy from linguistic complexity in the context of LCS patient information. We demonstrate that SLMs offer a safe, deployable, and resource-efficient alternative that maintains high fidelity to national guidelines while remaining accessible to low-literacy populations. Future work will focus on validating these findings in real-world patient cohorts to assess impact on screening uptake.

Article Details

Volume / Issue Vol. 44, Issue 16_suppl
Published June 01, 2026
ISSN 0732-183X
Publisher Lippincott Williams & Wilkins

Journal Info

Journal of Clinical Oncology

Lippincott Williams & Wilkins

ISSN: 0732-183X Health Sciences

Authors (5)

S

Sanjay Khanna

The Royal Marsden Hospital, London, United Kingdom

H

Hitesh Khanna

National Health Service, London, United Kingdom

Y

Yueqi Ge

The Royal Marsden Hospital, London, United Kingdom

G

Gina Sherpa

Imperial College Healthcare NHS Trust, London, United Kingdom

R

Richard Lee