Structured oncologist appraisal of AI-generated clinical literature summaries.
Abstract
e23296 Background: In alignment with the American Society of Clinical Oncology’s guiding principles for responsible use of AI in oncology, particularly accountability, transparency, and oversight, there is a need to understand how AI-generated clinical literature summaries perform under real-world physician review. Although AI tools are increasingly used to synthesize medical literature at the point of care, oncology poses unique challenges due to rapidly evolving evidence, complex guidelines, and comorbid conditions. The frequency and nature of AI-generated errors, and the role of physician oversight in identifying clinically ready outputs, remain incompletely characterized. Methods: We conducted an observational evaluation of AI-generated clinical literature summaries produced within an AI-assisted clinical reference platform (DoxGPT). Physicians performed structured clinical appraisal (PeerCheck), classifying summaries as acceptable for clinical application or rejected, with reasons documented for rejection. Time from outreach to review completion was assessed, as was review consistency among multiple reviewers. Data collection is ongoing, and descriptive statistics are presented. Results: To date, more than 30,000 physician appraisals have been completed across 416,829 AI-generated summaries. For summaries deemed clinically acceptable, the median number of reviewers was 2 (range, 1 to 26). Median time to review was 4 hours from email request. The largest specialty categories were Oncology (n = 39,326, 9.4%), Obstetrics and Gynecology (n = 22,816, 5.5%), and Cardiology (n = 17,511, 4.2%). Overall, 86% of summaries were accepted as ready for clinical application (n = 25,950). Rejection rates varied by specialty (1.8% to 31%), with higher rejection observed in medical genetics (31.3%, n = 96) and lower rejection observed in oncology and interventional radiology (1.8%, n = 59), noting limited sample sizes in some specialties. Among 4,216 rejected summaries, common reasons included missing data (38.8%, n = 1,635), factual errors (28.2%, n = 1,188), oversimplification (13.1%, n = 551), outdated information (8.4%, n = 353), citation errors (6.8%, n = 287), and misunderstanding the clinical question (1.7%, n = 70). Additional oncology-specific analyses will be presented. Conclusions: AI-based clinical reference tools can produce clinically useful outputs but remain prone to identifiable error types. Structured physician oversight is essential to ensure accuracy, completeness, and clinical applicability. In this evaluation, most AI-generated summaries were considered ready for clinical use, and physician review was completed rapidly, supporting the feasibility of timely human oversight. Missing clinically relevant information was the most common cause of rejection, underscoring the value of physician review in strengthening AI-assisted clinical information synthesis.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (6)
Kamal Menghrajani
1Memorial Sloan Kettering Cancer Center, New York, United States
James DeRosa
Doximity, San Francisco, CA
Mona Ascha
Doximity, San Francisco, CA
Anna Ransbotham-Cole
Doximity, San Francisco, CA
Arvind Rajan
Doximity, San Francisco, CA
Amit Phull
Doximity, San Francisco, CA