Clinical experience–based assessment of informed consent forms written by generative artificial intelligence.
Abstract
e13717 Background: The OpenEvidence (OE) generative artificial intelligence (AI) platform has emerged as a purpose-built tool for clinical use, with recent expansion into patient-based communication capabilities. The implementation of AI in generating informed consent documents has not yet been described for chemotherapy or immunotherapy. We sought to evaluate generalized systemic cancer treatment consent forms generated by OE and GPT 5.1 (OpenAI). Methods: Forms were generated using the same prompt for both OE and GPT 5.1. We developed an 8-item assessment instrument to evaluate forms based on description, accuracy, clarity, detail, and length. The assessment compared AI generated forms to a standard baseline consent form. Responses were collected on a 5-point scale ranging from -2 (worse) to +2 (better) with 0 representing equivalence. The instrument was administered through the Qualtrics platform, with raters randomized to evaluate one of the forms in a blinded format. Additionally, raters were blind to the AI generated nature of the forms. Raters were pharmacists, nurses, and physicians in the hematology/oncology department at our institution. Anonymized information regarding raters’ clinical practice background was collected. Composite instrument scores were calculated by summing item-level responses for each rater. Overall composite scores between OE and GPT 5.1 were compared using the Wilcoxon rank-sum test. Differences in composite scores across years-in-practice categories were evaluated separately for each model using the Kruskal–Wallis rank-sum test. Results: A total of 29 complete form assessments were collected: 16 for GPT 5.1 and 13 for OE. 79% (n = 23) of raters had an outpatient specialization. 59% (n = 17) of raters were nurses, 34% (n = 10) were pharmacists, and 7% (n = 2) were physicians. 34% (n = 10) of raters had < 10 years in practice and 32% (n = 9) raters had > 20 years in practice. Median (IQR) composite instrument scores for GPT 5.1 and OE were 6 (2, 14) and 8 (5, 10) respectively, p = 0.63. Scoring distribution did not differ significantly by rater training background (GPT 5.1: p = 0.79, OE: p = 0.97). Composite scores varied across years-in-practice groups for both models. For GPT 5.1, median (IQR) composite scores were 2 (−3, 5) for < 10 years, 13 (9, 16) for 10–20 years, and 11 (4, 15) for > 20 years (p = 0.07). For OE, median (IQR) composite scores were 9.0 (8.0, 11.0), 8.0 (5.8, 14.0), and 3.0 (−1.0, 4.0), respectively (p = 0.05). Conclusions: Despite the medically oriented focus of OE, informed consent documentation did not differ significantly between OE and the publicly available GPT 5.1. We observed variation in perceptions of these forms based on clinical experience, which may inform implementation of this technology.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (7)
Shiva Balasubramanian
1University of Louisville, Brown Cancer Center, Louisville, United States
Borna Amir-Kabirian
1University of Louisville, Brown Cancer Center, Louisville, United States
Tyler Jones
Shreyas Kalantri
University of Louisville, Louisville, KY
Katlyn Mulhall
Brown Cancer Center, University of Louisville, Louisville, KY
Maiying Kong
Goetz Hans Kloecker
Brown Cancer Center, University of Louisville, Louisville, KY