Graph informed biomarker discovery framework using transcriptomic machine learning for glioblastoma prognosis

O Osama Mahmoud M Mahmoud Mounir W Walaa Gad

Abstract

Abstract Identifying reproducible, interpretable prognostic signals from high-dimensional transcriptomics remains challenging because gene-level models often ignore network context. We developed Graph-Informed Biomarker Discovery (GIBD), a locked transcriptomics-only framework for primary glioblastoma that integrates RNA-seq expression with high-confidence STRING topology through weighted protein–protein interaction (WPPI) self-preserving feature construction. Model development, feature selection, scaler fitting, threshold selection, and locking used The Cancer Genome Atlas (TCGA) only, followed by post-lock external validation in the Chinese Glioma Genome Atlas (CGGA). The final TCGA cohort included 147 patients, and the empirical TCGA median overall survival of 357 days defined binary risk groups. The binary-evaluable CGGA cohort included 131 patients. The locked GIBD-XGBoost K100 model used 100 features (65 WPPI-derived, 35 raw-expression features) and threshold 0.53. TCGA out-of-fold AUC was 0.617. Post-lock CGGA validation yielded an AUC of 0.609, sensitivity of 73.9%, specificity of 50.6%, balanced accuracy of 62.3%, and a C-index of 0.537. SHAP identified TSPAN13 as the strongest global contributor, and full-transcriptome TCGA GSEA identified 156 terms at FDR < 0.05, with coherent high-risk inflammatory, hypoxic, metabolic, extracellular-matrix, complement/coagulation, and angiogenic enrichment. GIBD preserved an external transcriptomic risk-prioritization signal requiring prospective recalibration and multimodal validation before translational use.

Article Details

Volume / Issue Vol. 16, Issue 1
Published June 23, 2026
ISSN 2045-2322
Publisher Nature Portfolio

Journal Info

Scientific Reports

Nature Portfolio

ISSN: 2045-2322 Open Access Life Sciences

Authors (3)

O

Osama Mahmoud

M

Mahmoud Mounir

W

Walaa Gad