Pan-cancer gene set discovery via scRNA-seq for optimal deep learning based downstream tasks.

J Jong Hyun Kim (Department of Chemistry) Y Yeonuk Jeong (Lg AI Research, Seoul, South Korea) H Hj Yeo (Lg AI Research, Seoul, South Korea) Y Yong Min Park (Lg AI Research, Seoul, South Korea) H Hyung Kyung Kim (Samsung Medical Center, Seoul, South Korea) S Soonyoung Lee J Jongseong Jang

Abstract

e13606 Background: The application of deep learning to transcriptomics data has enabled various downstream tasks in cancer research. Due to the high dimensionality of transcriptomic data and inherent biological complexity, effective feature selection is crucial. Most feature selection methods rely on RNA sequencing data, which fails to capture the complex cellular heterogeneity within tumors. Single-cell RNA sequencing provides detailed insights into tumor heterogeneity and cellular dynamics masked in RNA-seq data. In this study, we demonstrate that gene sets derived from scRNA-seq data outperform RNA-seq-based selections in pan-cancer downstream tasks. Methods: We analyzed scRNA-seq data from 181 tumor biopsies across 13 cancer types and RNA-seq data from 7,178 TCGA tumor samples. High-dimensional weighted gene co-expression network analysis (hdWGCNA) was performed on the scRNA-seq data to identify co-expressed gene modules. Genes within these modules were further refined through XGBoost-based feature selection for specific downstream tasks. The selected gene sets were evaluated using two deep learning models (MLP and GNN) optimized for downstream tasks. Model performance was assessed on multiple downstream tasks using stratified cross-validation with 100 bootstrap iterations. Results: Through hdWGCNA, we identified 11 co-expression modules comprising 1,857 genes from scRNA data. These genes were refined using XGBoost and evaluated across 13 downstream tasks, compared against six published gene sets and OncoKB database. Our XGBoost-refined hdWGCNA gene sets achieved the highest performance in 12 tasks with MLP and 10 tasks with GNN, with minimal performance deviations in the remaining tasks. Using the MLP model, we achieved high performance across all tasks. For MSI classification in STAD and CRC, our model achieved AUROCs of 0.990 and 0.931, respectively. TMB assessment in CRC, LUAD, SKCM, and LUSC showed AUROCs of 0.826, 0.791, 0.772, and 0.647, respectively. In mutation prediction, AUROCs were 0.904, 0.869, 0.868, 0.845, and 0.790 for PAAD-KRAS, LUAD-TP53, LUAD-EGFR, LUAD-KRAS, and STAD-TP53, respectively. Feature importance analysis identified DPM1 as significant across all tasks, while BAD and FKBP4 showed high importance in 12 and 10 tasks respectively, suggesting their potential as pan-cancer biomarkers. Conclusions: This study proposes a comprehensive framework integrating scRNA-seq data with advanced features selection methods for pan-cancer analysis. Our XGBoost-refined hdWGCNA gene sets demonstrated superior performance across multiple downstream tasks. Notably, we identified DPM1, BAD, and FKBP4 as potential pan-cancer biomarkers, which consistently showed significance across various cancer types. These findings present a refined approach for enhancing predictive modeling in cancer genomics and offer promising targets for future therapeutic development.

Article Details

Volume / Issue Vol. 43, Issue 16_suppl
Published June 01, 2025
ISSN 0732-183X
Publisher Lippincott Williams & Wilkins

Journal Info

Journal of Clinical Oncology

Lippincott Williams & Wilkins

ISSN: 0732-183X Health Sciences

Authors (7)

J

Jong Hyun Kim

Department of Chemistry

Y

Yeonuk Jeong

Lg AI Research, Seoul, South Korea

H

Hj Yeo

Lg AI Research, Seoul, South Korea

Y

Yong Min Park

Lg AI Research, Seoul, South Korea

H

Hyung Kyung Kim

Samsung Medical Center, Seoul, South Korea

S

Soonyoung Lee

J

Jongseong Jang