Comparative analysis of large language models (LLMs) for pharmacogenomic (PGx)–guided prescribing in oncology.
Abstract
e13665 Background: PGx is increasingly applied in oncology to personalize drug therapy, minimize toxicity, and improve efficacy. LLMs have shown potential in providing medical recommendations, though their performance in PGx is not well understood. This study evaluates LLMs’ performance in providing chemotherapy prescribing recommendations based on PGx results using a modified version of the previously developed Generative AI Performance scoring (G-PS) tool. Methods: Six single-gene and four multi-gene oncology PGx theoretical cases were developed for gene-drug pairs with established prescribing guidance: DPYD-fluoropyrimidines, UGT1A1-irinotecan, and TPMT/NUDT15-thiopurines. Prompts followed the structure: “ My patient is a [gene] [phenotype] with a genotype of [genotype]. How does this result affect [drug] prescribin g?” . LLMs tested included ChatGPT 5.2, Copilot (Think Deeper), Gemini (Google AI Overview), and OpenEvidence. Each case was queried three times per LLM on a private browser, except OpenEvidence, which was queried once due to login requirement. Each response was evaluated for accuracy per guidelines (FDA, CPIC, or NCCN) and weighted across up to three domains: drug selection, dose selection, and dose modification. Penalty was applied per the number and severity of hallucinations. A modified G-PS (mG-PS) grading scheme generated a final score for each LLM response, ranging from -1 (all hallucinations) to 1 (fully correct). Grading was performed by oncology and PGx pharmacists and a pharmacy student. Results: Table 1 reports mean mG-PS and percentage of responses with at least one hallucination or missing information across all LLMs. Across 10 cases, OpenEvidence had the highest mean mG-PS (0.88) and the lowest rate of hallucination (10%). Among LLMs tested in triplicate, ChatGPT had the highest mean mG-PS (0.82), followed by Copilot (0.68), and Gemini (0.54). ChatGPT and Copilot had the highest rates of hallucination (48%), followed by Gemini (35%). Gemini most frequently provided inadequate or no recommendations (60%). Conclusions: Performance across LLMs was variable. OpenEvidence showed the highest accuracy with the fewest hallucinations, though its use remains limited to healthcare professionals. Among publicly available LLMs, ChatGPT performed comparatively well but still produced high rates of hallucinations and incomplete information. Clinicians should be cautious when using LLMs in clinical practice, and in the case of PGx, continue to rely on guidelines such as CPIC. The modified G-PS tool offers a structured framework for assessing LLM performance in PGx and can be applied to more complex multi-gene PGx cases in oncology and supportive care. ChatGPT Copilot Gemini OpenEvidence Mean mG-PS 0.82 0.68 0.54 0.88 Any Hallucinations 48% 48% 35% 10% Major Hallucinations 7% 13% 10% 10% Minor Hallucinations 42% 37% 25% 0% Missing Information 37% 45% 60% 30%
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (7)
Mai Nguyen
Grace Nguyen
Creighton University School of Medicine (Phoenix Regional Campus), Phoenix, AZ
Sarah Morris
Atrium Health Levine Cancer Institute, Charlotte, NC
Natalie Marie Reizine
University of Illinois College of Medicine at Chicago, Division of Hematology and Oncology, Chicago, IL
Jai Narendra Patel
Atrium Health Levine Cancer Institute, Charlotte, NC
Ryan Huu-Tuan Nguyen
University of Illinois College of Medicine at Chicago, Division of Hematology and Oncology, Chicago, IL
Noor Naffakh
University of Illinois at Chicago, Chicago, IL