PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 20, 2026Briefings in Bioinformatics0 citationsOpen Access

Rank-based learning: a novel high-throughput algorithm resilient to missing data and effective for datasets with small sample size

View Full Paper
LSLulu SongHRHamid Khoshfekr RudsariJFJohannes F Fahrmann

Key Points

  • The aim is to develop a robust Rank-Based Learning method for binary classification in high-throughput omics data, particularly with small sample sizes.
  • Introduced a Rank-Based Learning (RBL) algorithm focused on feature rankings
  • Compared RBL against Logistic Regression (LR) and Random Forest (RF)
  • Evaluated using simulated data and two real-world proteomics datasets
  • Assessed performance using Area Under the Curve (AUC) metrics
  • RBL outperformed LR and RF in simulations with batch effects and missing data
  • Achieved a test AUC of 0.76 in small cell lung cancer (SCLC)
  • In duodenopancreatic neuroendocrine tumors (dpNET), RBL reached an AUC of 0.83 on the development set
  • RBL's focus on relative feature rankings reduces the effect of non-biological variability

Abstract

Abstract High-throughput omics data present challenges for binary classification due to platform variability, batch effects, missing values, and high dimensionality. This study presents a novel Rank-Based Learning (RBL) method that leverages relative feature rankings to improve robustness and generalizability. We evaluated RBL against established methods like Logistic Regression (LR) and Random Forest (RF) using simulated data and two real-world plasma proteomics datasets: early-stage small cell lung cancer (SCLC) and duodenopancreatic neuroendocrine tumors (dpNET) in patients with Multiple Endocrine Neoplasia type 1 (MEN1). In simulation experiments, RBL outperformed LR under conditions involving batch effects, missing data, and varying numbers of true differential features. In SCLC, RBL yielded a test AUC of 0.76 (95% CI: 0.42–1.00), surpassing LR with Lasso (0.65 95% CI: 0.47–0.84) and RF with feature importance (0.59 95% CI: 0.33–0.87). In dpNET, RBL achieved an AUC of 0.83 (95% CI: 0.67–0.97) on the development set and 0.80 (95% CI: 0.54–0.98) on the test set, outperforming LR with Lasso (0.57 95% CI: 0.40–0.77) and RF with feature importance (0.53 95% CI: 0.29–0.77). By emphasizing feature ranking rather than absolute expression levels, RBL effectively mitigates the impact of non-biological variation. Overall, RBL improves the predictive accuracy of diagnostic models for complex diseases and provides a promising framework for developing more reliable and generalizable diagnostic tools from omics data, moving them closer to clinical application.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Song et al. (2025) studied this question.

synapsesocial.com/papers/6997fa80ad1d9b11b3453cdahttps://doi.org/10.1093/bib/bbaf666
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Classifying breast cancer subtypes on multi-omics data via sparse canonical correlation analysis and deep learning2024 · 46 citations
  2. 2Three Grand Challenges in High Throughput Omics Technologies2022 · 8 citations
  3. 3Multi-omic data integration and analysis using systems genomics approaches: methods and applications in animal production, health and welfare2016 · 215 citations
  4. 4From Omics to Multi-Omics: A Review of Advantages and Tradeoffs2024 · 86 citations
  5. 5Identification of Transferrin Receptor 1 (TfR1) Overexpressed in Lung Cancer Cells, and Internalization of Magnetic Au-CoFe2O4 Core-Shell Nanoparticles Functionalized with Its Ligand in a Cellular Model of Small Cell Lung Cancer (SCLC)2022 · 25 citations