Application of unsupervised genetic clustering to identify biologically distinct cholangiocarcinoma subtypes with differential survival.
Abstract
4025 Background: Cholangiocarcinoma (CCA) is a genetically heterogeneous malignancy for which current anatomic and histologic classifications inadequately predict survival. We hypothesized that unsupervised machine learning clustering of mutations and copy number alterations(CNA) could identify biologically distinct CCA subgroups with clinically meaningful differences in survival. Methods: Genomic and clinical data were obtained from the MSK cholangiocarcinoma dataset from cBioPortal. 790 patients with intrahepatic and extrahepatic CCA with available mutation and CNA data were included. Multiple unsupervised approaches were evaluated, including k-means, hierarchical clustering, non-negative matrix factorization, PCA–k-means, and UMAP–k-means. Model performance was assessed using silhouette scores, with UMAP–k-means (n=4 clusters) selected as the optimal method. Clusters were defined by dominant genetic alterations and compared for overall survival using Kaplan–Meier analysis with pairwise log-rank testing. Subgroup analyses included patients without curative-intent surgery and a young-onset cohort. Results: UMAP–k-means delineated four distinct genomic clusters ordered by progressively worse median survival. Cluster 1, defined by ARID1A or BAP1 mutations, demonstrated the most favorable survival outcomes. Cluster 2 consisted of tumors wild type for recurrent driver alterations. Cluster 3 was characterized by TP53 , KRAS , or IDH1 mutations. Cluster 4, defined by CDKN2A double deletion or ERBB2 amplification, exhibited the poorest median survival. Overall survival differed significantly across clusters, with all pairwise Kaplan–Meier comparisons reaching statistical significance except between clusters 1 and 2. These survival differences remained significant in patients who did not undergo curative-intent surgery. Notably, cluster 4 remained associated with significantly worse survival within the young-onset cohort. Conclusions: Unsupervised genomic clustering using UMAP–k-means identified biologically distinct subtypes of cholangiocarcinoma, with clinically significant survival differences that persist across surgical and age-based subgroups. These findings support future genomic subtyping as a prognostic framework and a rationale for biology-driven clinical trial stratification in cholangiocarcinoma. Unsupervised genetic clustering identifies biologically distinct cholangiocarcinoma subtypes with differential survival. Cluster Primary Genetic alterations N patients Median OS log-rank test P-value vs cluster 2 log-rank test P-value vs cluster 3 log-rank test P-value vs cluster 4 1 ARID1A or BAP1 mutant 120 28.0 0.308 <0.001 <0.001 2 Wild types 224 26.7 - 0.003 <0.001 3 TP53, KRAS or IDH1 mutant 319 22.0 - 0.002 4 CDKN2A deletion or ERBB2 amplification 127 16.9 -
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (12)
Qianchen Zhang
Carle Illinois College of Medicine, Urbana, IL
Fumihiro Kawano
Carle Foundation Hospital, Urbana, IL
Daniel Sing Han Cheah
Carle Illinois College of Medicine, Urbana, IL
Kathryn Chen Tsai
Carle Illinois College of Medicine, Urbana, IL
Helen Kemprecos
Carle Illinois College of Medicine, Urbana, IL
Alshammary Shadi
Carle Foundation Hospital, Urbana, IL
Arundhati Pillai
Carle Illinois College of Medicine, Urbana, IL
Megha Vijay Guggari
Carle Illinois College of Medicine, Urbana, IL
Gregory Polites
Carle Foundation Hospital, Urbana, IL
Mark Cohen
Carle Illinois College of Medicine, Urbana, IL
Zeynep Madak Erdogan
University of Illinois Urbana-Champaign, Urbana, IL
Claudius Conrad
Carle Illinois College of Medicine, Urbana, IL