Comparative performance of large language model derived and clinician documented ECOG-PS in predicting clinical outcomes in head and neck squamous cell carcinoma (HNSCC) undergoing concurrent chemoradiation (CRT).
Abstract
e23101 Background: Clinician-documented ECOG-PS (dECOG-PS) often overestimates true ECOG-PS. The ability of large language models (LLMs) to infer functional status from clinical descriptors remains limited in non-English texts and with open-weight models suitable for low-resource deployment. Methods: Clinical notes from the visit immediately preceding CRT for HNSCC pts treated between 2011 and 2021 were identified. We optimized a zero-shot structured reasoning prompt with examples and applied it to six locally deployed LLMs: Gemma3-27B (GEM), GPT-OSS-20B (GPT), Ministral-3-14B (MIN), Mistral-Small-24B (MIS), DeepSeek-R1-14B (DPS), and Qwen3-14B (QWE) to infer ECOG-PS (llmECOG-PS), without outcome information. We then assessed llmECOG performance against dECOG-PS for five outcomes: incomplete CRT (≤200 mg/m² cumulative cisplatin), emergency department visits within 7 days (ER7D) and 30 days (ER30D), and mortality at 30 days (M30D) and 180 days (M180D). Discrimination was assessed using AUC, and effect sizes estimated by logistic regression odds ratios (OR) for ECOG-PS ≥2 versus 0–1. Results: Of 693 pts, dECOG was 0 (17.5%), 1 (73.9%), 2 (8.2%), and 3 (0.3%). Overall agreement (llmECOG = dECOG), down-classification (llmECOG < dECOG), and up-classification (llmECOG > dECOG) rates were: GPT 53.8%, 22.9%, 23.2%; MIS 51.8%, 4.1%, 44.1%; MIN 57.4%, 3.4%, 39.3%; GEM 65.6%, 5.3%, 29.1%; DPS 58.1%, 5.3%, 36.6%; and QWE 65.4%, 8.8%, 25.7%. MIN and GPT llmECOG showed superior discrimination for M30D (AUC: MIN 0.775, GPT 0.768 vs dECOG 0.651); however, only MIN outperformed dECOG for M180D (AUC 0.609 vs 0.594). All models outperformed dECOG for ER7D (AUC range 0.553–0.602 vs dECOG 0.549) and ER30D (AUC range 0.547–0.574 vs dECOG 0.546), except MIS and GEM. No model improved discrimination for incomplete CRT. Effect sizes for grouped ECOG (≥2 vs 0–1) were often higher and more frequently significant (OR > 1) for llmECOG than for dECOG (Table 1). Conclusions: Smaller, open-weight, privacy-focused model-derived ECOG showed improved discrimination compared to dECOG for emergency visits and mortality after CRT. This indicator may be useful for quality-of-care monitoring and risk stratification, particularly in low-resource settings. GPT MIS GEM MIN DPS QWE dECOG Incomplete CRT 1.6 (1.0-2.5) 1.5 (1.0-2.1) 1.5 (0.9-2.3) 1.3 (0.9-1.9) 1.3 (0.8-1.9) 2.1 (1.3-3.3) 2.3 (1.3-4.0) ER7D 2.4 (1.5-3.9) 2.2 (1.5-3.4) 1.5 (0.9-2.4) 1.7 (1.1-2.6) 1.9 (1.2-2.9) 2.3 (1.4-3.7) 2.0 (1.1-3.7) ER30D 1.8 (1.2-2.7) 1.9 (1.4-2.6) 1.5 (1.0-2.3) 1.7 (1.2-2.4) 1.7 (1.2-2.4) 1.8 (1.2-2.8) 1.9 (1.1-3.2) M30D 7.5 (1.2-57.2) 2.7 (0.4-20.9) 3.4 (0.4-20.1) 9.1 (1.3-178.7) 0.7 (0.0-4.5) 3.9 (0.5-24.0) 3.6 (0.2-28.8) M180D 1.8 (0.9-3.4) 1.6 (0.9-2.9) 2.1 (1.1-4.0) 2.4 (1.3-4.3) 1.5 (0.8-2.8) 2.4 (1.2-4.6) 2.3 (1.0-5.0)
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (8)
Gabriel Berlingieri Polho
Instituto do Câncer do Estado de São Paulo, São Paulo, Brazil
João Pedro Aparecido Ribas De Souza
Instituto do Câncer do Estado de São Paulo, São Paulo, Brazil
Rafael Lara Nohmi
University of São Paulo, São Paulo, Brazil
Amanda Acioli A. Robatto
Instituto do Câncer do Estado de São Paulo, São Paulo, Brazil
Luciana Carvalho Barros
Instituto do Câncer do Estado de São Paulo, São Paulo, Brazil
Milena Perez Mak
Instituto do Cancer do Estado de São Paulo - Faculdade de Medicina da Universidade de Sao Paulo, Sao Paulo, Brazil
Gilberto Castro
Instituto do Câncer do Estado de São Paulo, São Paulo, Brazil
Felippe Lazar Neto
The University of Texas MD Anderson Cancer Center, Houston, TX