Pan-cancer gene set discovery via scRNA-seq for optimal deep learning based downstream tasks.
Abstract
e13606 Background: The application of deep learning to transcriptomics data has enabled various downstream tasks in cancer research. Due to the high dimensionality of transcriptomic data and inherent biological complexity, effective feature selection is crucial. Most feature selection methods rely on RNA sequencing data, which fails to capture the complex cellular heterogeneity within tumors. Single-cell RNA sequencing provides detailed insights into tumor heterogeneity and cellular dynamics masked in RNA-seq data. In this study, we demonstrate that gene sets derived from scRNA-seq data outperform RNA-seq-based selections in pan-cancer downstream tasks. Methods: We analyzed scRNA-seq data from 181 tumor biopsies across 13 cancer types and RNA-seq data from 7,178 TCGA tumor samples. High-dimensional weighted gene co-expression network analysis (hdWGCNA) was performed on the scRNA-seq data to identify co-expressed gene modules. Genes within these modules were further refined through XGBoost-based feature selection for specific downstream tasks. The selected gene sets were evaluated using two deep learning models (MLP and GNN) optimized for downstream tasks. Model performance was assessed on multiple downstream tasks using stratified cross-validation with 100 bootstrap iterations. Results: Through hdWGCNA, we identified 11 co-expression modules comprising 1,857 genes from scRNA data. These genes were refined using XGBoost and evaluated across 13 downstream tasks, compared against six published gene sets and OncoKB database. Our XGBoost-refined hdWGCNA gene sets achieved the highest performance in 12 tasks with MLP and 10 tasks with GNN, with minimal performance deviations in the remaining tasks. Using the MLP model, we achieved high performance across all tasks. For MSI classification in STAD and CRC, our model achieved AUROCs of 0.990 and 0.931, respectively. TMB assessment in CRC, LUAD, SKCM, and LUSC showed AUROCs of 0.826, 0.791, 0.772, and 0.647, respectively. In mutation prediction, AUROCs were 0.904, 0.869, 0.868, 0.845, and 0.790 for PAAD-KRAS, LUAD-TP53, LUAD-EGFR, LUAD-KRAS, and STAD-TP53, respectively. Feature importance analysis identified DPM1 as significant across all tasks, while BAD and FKBP4 showed high importance in 12 and 10 tasks respectively, suggesting their potential as pan-cancer biomarkers. Conclusions: This study proposes a comprehensive framework integrating scRNA-seq data with advanced features selection methods for pan-cancer analysis. Our XGBoost-refined hdWGCNA gene sets demonstrated superior performance across multiple downstream tasks. Notably, we identified DPM1, BAD, and FKBP4 as potential pan-cancer biomarkers, which consistently showed significance across various cancer types. These findings present a refined approach for enhancing predictive modeling in cancer genomics and offer promising targets for future therapeutic development.
Article Details
Journal Info
Journal of Clinical Oncology
Lippincott Williams & Wilkins
Authors (7)
Jong Hyun Kim
Department of Chemistry
Yeonuk Jeong
Lg AI Research, Seoul, South Korea
Hj Yeo
Lg AI Research, Seoul, South Korea
Yong Min Park
Lg AI Research, Seoul, South Korea
Hyung Kyung Kim
Samsung Medical Center, Seoul, South Korea
Soonyoung Lee
Jongseong Jang