Evaluating AI-driven auditing of mammography reports for BI-RADS compliance.
Abstract
e13720 Background: The Breast Imaging Reporting and Data System (BI-RADS) provides standardized guidelines for accurate reporting of mammograms which is essential for breast cancer screening and diagnosis. This study evaluates the potential of artificial intelligence (AI) systems as tools for auditing mammography reports. Methods: Medical records of 600 patients were reviewed and patients with reported diagnostic &/or screening mammograms were included in the study. 528 anonymized mammography reports were analyzed for compliance with BI-RADS using four AI engines generating 2,112 runs. Each component of a mammogram report is assigned a percentage weight reflecting its clinical relevance. The assigned weight of each component is calculated and designated as complete, partially complete, or missing. Then, the compliance level of each individual report based on the cumulative score of the components were categorized into complete, somehow complete, satisfactory, deficient or severely deficient. Fleiss' kappa & Cohen's κ assessed inter-rater reliability. Results: 528 female patients included with mean age of 53 years + 12 years, 361 (68.40%) patients were married, (48.20%) were obese and ( 29.7%) were overweight. Auditing mammography reports with different AI tools showed high overall compliance scores (4.24–4.31/5), with compliance category score per AI tools listed in Table 1. Moderate inter-system overall agreement (Fleiss' kappa = 0.567, p < .001) was observed. Cohen's κ was run individually to determine if different AI engines agreed on compliance levels; GPT showed moderate agreement with Co-pilot (κ = 0.574 (95% CI, 0.515 to 0.633), p < .001), Perplexity (κ = 0.576 (95% CI, 0.517 to 0.635), p < .001), and Claude (κ = 0.528 (95% CI, 0.467 to 0.589), p < .001). Co-Pilot showed moderate agreement with Perplexity (κ = 0.504 (95% CI, 0.441 to 0.567), p < .001), and Claude (κ = 0.515 (95% CI, 0.452 to 0.578), p < .001). Perplexity showed substantial agreement with Claude (κ = 0.705 (95% CI, 0.654 to 0.756), p < .001). Non-breast cancer cases had higher compliance scores than biopsy-proven cancer cases. No significant correlations were observed between compliance and age, BMI, or marital status. Conclusions: AI systems are promising tools for assessing mammography reporting quality. They can analyze reports instantly and be integrated automatically to reporting template to ensure adherence and comprehensiveness. A dedicated AI models for radiological reports interpretation can enhance the quality of reporting. Compliance category scores as per individual AI tool. AI Engine Complete Somehow complete Satisfactory Deficient Severely deficient GPT 246 (46.60%) 189 (35.80%) 74 (14.00%) 13 (2.50%) 6 (1.10%) Co-Pilot 265 (50.19%) 194 (36.74%) 45 (8.52%) 17 (3.22%) 7 (1.33%) Perplexity 255 (48.28%) 190 (36.00%) 60 (11.40%) 17 (3.22%) 6 (1.10%) Claude 253 (48.00%) 180 (34.10%) 73 (13.80%) 16 (3.00) 6 (1.10%)
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (6)
Atlal M. Abusanad
King Abdulaziz University, Jeddah, Jeddah, Saudi Arabia
Abdallah Bokhary
Faculty of Medicine, King Abdulaziz University, Jeddah, Saudi Arabia
Abdulrahman Almaghrabi
Faculty of Medicine, King Abdulaziz University, Jeddah, Saudi Arabia
Ahmed Qumsani
Faculty of Medicine, King Abdulaziz University, Jeddah, Saudi Arabia
Raad Alsaadi
Faculty of Medicine, King Abdulaziz University, Jeddah, Saudi Arabia
Khalid Basahih
Faculty of Medicine, King Abdulaziz University, Jeddah, Saudi Arabia