Fine-scaled genomic ancestry clusters to reveal population-specific cancer enrichments in the All of Us Research Program.
Abstract
10514 Background: Specific populations exhibit known enrichments for certain cancers, such as breast cancer (BC) in Ashkenazi Jewish individuals. However, identification typically relies on self-reported race or broad continental genomic labels, which can obscure risks in understudied or admixed U.S. populations. Methods: Using All of Us (AOU) v8 (n=415k), participants with electronic health records (EHR) and whole-genome sequencing (WGS) were grouped into clusters based on genomic segments identical by descent. Clusters were labeled using genomic and anthropological references. We curated clinical data to remove duplicate or mislabeled entries (e.g., misidentifying metastatic lesions as primary tumors). Enrichment was determined via logistic regression, adjusting for age, sex, BMI, socioeconomic status, smoking, and recruitment site, with correction for multiple hypothesis testing using Benjamini-Hochberg. Results: A total of 245k individuals with EHR and quality-controlled WGS were assigned to 58 clusters; 41876 individuals were identified with at least one primary cancer diagnosis, spanning 19 broad types. A non-comprehensive selection of model results are presented in Table 1. These include previously established risks, such as BC enrichment in Ashkenazi-Jewish populations and prostate cancer in African Americans compared to Non-Jewish Europeans (NJE). Other, more novel findings include significant BC enrichment in admixed Hawaiians and colorectal cancer (CRC) in Puerto Rican and Dominican groups (not observed in the African American cluster). While individuals of Mexican ancestry showed lower lung cancer enrichment, this was not consistent in other admixed Hispanic populations, such as Colombians. Conclusions: Fine-scaled genomic clustering identifies cancer risks with higher precision than race-based or broad ancestral statistics. This methodology facilitates the identification of globally rare, potentially pathogenic variants that are enriched within specific subpopulations. These results underscore the necessity of granular ancestry data in understanding the genetic architecture of complex diseases like cancer. Cancer Population Population Case N (%) NJE Case N (%) Odds Ratio (95% CI) q-val BC 1 Ashkenazi Jewish 915 (10.6%) 6.9% 1.33 (1.23-1.44) <0.001 Prostate 2 African American 852 (5.4%) 8.0% 1.51 (1.38-1.65) <0.001 BC 1 Hawaiian <20* (9.4%) 6.9% 2.64 (1.39-5.03) 0.040 CRC Puerto Rican 89 (1.5%) 1.5% 1.54 (1.24-1.93) 0.003 CRC Dominican 81 (1.4%) 1.5% 1.85 (1.32-2.61) 0.008 CRC African American 491 (1.2%) 1.5% 1.16 (1.04-1.30) 0.06 Lung Mexican 95 (0.4%) 1.1% 0.66 (0.53-0.82) 0.005 Lung Colombian <20* (0.7%) 1.1% 0.92 (0.33-2.54) 0.97 1: Percentages only include female participants. 2: Percentages only include male participants. *: Counts of less than 20 are hidden to comply with AoU reporting requirements.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (9)
Hersh Gupta
Mariko Isshiki
Department of Genetics, Albert Einstein College of Medicine, Bronx, NY
Defne Ercelen
Department of Genetics, Albert Einstein College of Medicine, Bronx, NY
Sharon Liu
Department of Genetics, Albert Einstein College of Medicine, Bronx, NY
Dana Luong
Department of Genetics, Albert Einstein College of Medicine, Bronx, NY
Chynna Smith
Department of Genetics, Albert Einstein College of Medicine, Bronx, NY
David Yang
John Greally
Department of Genetics, Albert Einstein College of Medicine, Bronx, NY
Srilakshmi M. Raj
Department of Genetics, Albert Einstein College of Medicine, Bronx, NY