Mutation rate differences across populations and association with performance disparities in pathology AI diagnostic models.
Abstract
1603 Background: Previous studies have established artificial intelligence (AI) algorithms to classify cancer types, providing real-time diagnostic support. In addition, AI models have identified previously unknown pathology patterns associated with cancer genomic profiles. However, these models exhibit variable performance in different demographic groups, and the causes remain largely unknown. To address this challenge, we investigated the relationships between biases in AI diagnostic models and mutation rate disparities across populations and evaluated the efficacy of a fairness-aware contrastive learning (FACL) framework in reducing performance disparities. Methods: We obtained whole-slide pathology images, mutation rates of the 5 most frequently mutated genes in each cancer type, age, sex, and race from 9,217 patients in The Cancer Genome Atlas across 10 cancer types. We identified tasks with performance disparities across demographic groups, and employed generalized linear models to quantify the relationship between mutation rates and model bias in each cancer type. We further developed an FACL framework, and evaluated its effectiveness in mitigating these disparities using metrics including differences in accuracy (DIA) and equal opportunity. Results: Six genomic profile prediction tasks showed significant performance disparities across population groups (Table). Variations in TP53 mutation are associated with differential error rates in serous UCEC v. nonserous UCEC, mixed IDC v. ILC, LUAD v. LUSC, and GBM v. LGG classification tasks. Differences in CDH1 mutation rates were linked to racial disparity in mixed IDC v. ILC and age discrepancy in IDC v. ILC classification tasks. Our FACL framework mitigated performance disparities across demographic groups in 5 out of 6 tasks where standard AI model exhibited significant bias (p < 0.05). Conclusions: Biases in AI-driven cancer pathology diagnosis stem from disparities in somatic mutation prevalence across demographic groups. Addressing these biases is critical to ensuring fairness and the global applicability of AI tools. Our findings demonstrate that the FACL-based framework effectively reduces performance disparities, making AI-powered cancer diagnostics more reliable. Tasks Mutation Sensitive Attribute Groups and Mutation rates Standard (S) v. FACL (F) models DIA sUCEC v. nsUCEC TP53 Race W 0.34 p<0.001 S: 0.13±0.10, p<0.001F: 0.10±0.05, p=0.088 B 0.45 Mixed IDC v. ILC CDH1 Race W 0.29 p<0.001 S: 0.07±0.02, p<0.001F: 0.11±0.04, p=0.233 B 0.46 TP53 Race W 0.15 p=0.047 S: 0.12±0.02, p=0.023F: 0.10±0.05, p=0.196 A 0.11 IDC v. ILC CDH1 Age ≥59 yrs 0.11 p=0.038 S: 0.05±0.02, p=0.001F: 0.05±0.05, p=0.370 <59 yrs 0.17 LUAD v. LUSC TP53 Sex F 0.75 p=0.002 S: 0.12±0.02, p<0.001F: 0.01±0.01, p=0.154 M 0.58 GBM v. LGG TP53 Race W 0.50 p=0.021 S: 0.20±0.01, p<0.001F: 0.29±0.02, p=0.005 B 0.33 W: White; B: Black; A: Asian.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (15)
Po-Jen Lin
Department of Biomedical Informatics, Harvard Medical School, Boston, MA
Shih-Yen Lin
Pei-Chen Tsai
Department of Biomedical Informatics, Harvard Medical School, Boston, MA
Fang-Yi Su
Department of Biomedical Informatics, Harvard Medical School, Boston, MA
Chun-Yen Chen
Fuchen Li
Department of Biomedical Informatics, Harvard Medical School, Boston, MA
Junhan Zhao
Yuk Yeung Ho
Department of Biomedical Informatics, Harvard Medical School, Boston, MA
Tsung-Lu Michael Lee
Department of Computer Science and Information Engineering, Southern Taiwan University of Science and Technology, Tainan, Taiwan
Elizabeth Healey
Ting-Wan Kao
Irene Tai-Lin Lee
Emory University, Atlanta, GA
Eric Chongze Ma
Western Connecticut Medical Group, PC., Danbury, CT
Jung-Hsien Chiang
Kun-Hsing Yu