Using artificial intelligence (AI) and retrieval-augmented generation (RAG) to identify patient-reported symptoms post-ambulatory surgery for cancer.
Abstract
e18003 Background: As ambulatory procedures become commonplace in oncology, post-surgical remote symptom monitoring (RSM) is important for timely responses to emergent problems. Patient-reported outcomes (PROs) are the gold standard for the capture of the patient symptomatic experience; however, existing electronic PRO (ePRO) assessments that include free text reporting are not well-equipped to capture emergent symptoms or nuanced concerns due to resource constraints. Recent significant advances in AI modeling may allow for a minimally burdensome opportunity to monitor patients’ recovery. As a proof of concept, we aimed to determine whether large language models (LLMs) could efficiently classify verbatim free text patient-reported symptoms post-ambulatory surgery. Methods: Secure enterprise cloud-based versions of GPT-5.2 and Claude Sonnet 4.5, licensed to our institution, were used to identify symptoms from verbatim free-text responses collected from a retrospective sample of N = 1,070 cancer patients post-ambulatory surgical procedures for up to 10 days. Responses from thyroidectomy patients (n = 147) were coded and verified by a surgeon into 161 unique symptoms (e.g., ‘headache’, ‘dizziness’, ‘swollen incision’) and used as the gold standard. Coded responses were incorporated into RAG for postulated performance boost. The fine-tuned model was then used to analyze responses from mastectomy (n = 252) and prostatectomy (n = 671) patients. Results: GPT-5.2 with RAG identified 360 symptoms post-thyroidectomy, including (ordered by prevalence) ‘headache’ (area under the receiver operating characteristic curve [AUC]=0.89, 95% Confidence Interval [CI]: 0.83, 0.94), ‘dizziness’ (AUC=0.72, 95% CI: 0.72), ‘cough’ (AUC=0.81, 95% CI: 0.73, 0.90), ‘cough’ (AUC=0.99, 95% CI: 0.99, 1.00), and ‘sore throat’ (AUC=0.78, 95% CI: 0.65, 0.88). Patients differed in words chosen (e.g., ‘pain on swallowing’, ‘burning with swallowing’, and ‘difficulty swallowing’). Semantic similarity incorporated into modeling yielded a taxonomy of 98 symptom categories and an overall performance at 0.93 precision, 0.92 recall, and 0.93 F1 score. Claude Sonnet 4.5 had comparable results. Removing RAG resulted in a negligible reduction in model performance. In mastectomy and prostatectomy respectively, the most prevalent symptoms were ‘swelling’, ‘pain’, ‘drainage’, and ‘numbness’; and ‘swelling’, ‘pain’, ‘hematuria’, and ‘bleeding.’ Conclusions: Real-time processing of patient-reported free text post-ambulatory surgical concerns may be feasible using AI. RAG did not confer an advantage, due in part to the limited number of human-annotated symptoms. If further validated, AI-assisted RSM can offer patients an efficient and unfiltered pathway for reporting their post-surgical experiences in a way that can be integrated into clinical decision-making.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (4)
Jennifer R. Cracchiolo
Memorial Sloan Kettering Cancer Center, New York, NY
Thomas M. Atkinson
Memorial Sloan Kettering Cancer Center, New York, NY
Aleksandr Petrov
Memorial Sloan Kettering Cancer Center, New York, NY
Yuelin Li
Guangdong Provincial Key Laboratory of Applied Botany, South China Botanical Garden, Guangzhou, China.