Gliomas represent the most prevalent type of brain tumor, with their most aggressive variant, glioblastoma multiforme, associated with high mortality rates. Due to their elevated molecular heterogeneity, accurate classification of gliomas has presented significant challenges. Therefore, considerable effort has been dedicated to identifying relevant biomarkers that improve early diagnosis and unveil new areas for treatment. Advances in high-throughput sequencing technology have enabled public resources such as The Cancer Genome Atlas (TCGA) to provide large-scale data from various cancers, allowing researchers to perform more comprehensive analysis of this disease. In this study, we introduce MOHVAE-B, a comprehensive framework designed for the integration of multi-omics data and biomarker discovery using data from TCGA. MOHVAE-B employs a supervised hierarchical variational autoencoder integrated with SHAP-based interpretability to effectively integrate high-dimensional multi-omics data and extract the most influential features driving the model’s predictions. Subsequently, Bayesian Networks (BNs) are constructed to model conditional dependencies between the selected features, providing insights into their possible relations. Applied to the TCGA glioma cohorts, MOHVAE-B achieved a near-perfect AUC of 0.9993 and successfully identified high-impact features related to glioma classification. For glioblastoma multiforme, this included six novel candidates: LINC02172, NACA2, LINC01114, HNRNPA1P48, PPIAL4G, and LINC01558. For low-grade gliomas, the model highlighted AMER2 as a promising marker. Across both cohorts, PMP2 stood out as a particularly strong candidate for a potential role in glioma pathogenesis. The constructed BNs provided an additional layer of validation, reinforcing NACA2 as a candidate of interest in glioma biology.
Silva et al. (Mon,) studied this question.