PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 10, 2026Scientific Reports0 citationsOpen Access

Ensemble machine learning algorithms leveraged on serum proteomics for enhanced early detection of ovarian cancer

JDJie DingXDXin DongLCLi Cao

Key Points

  • This study aims to enhance early detection of ovarian cancer through serum proteomics and machine learning.
  • Analyzed serum samples from 188 ovarian cancer patients and 208 healthy controls using MALDI-TOF mass spectrometry.
  • Developed an ensemble pipeline using eight machine learning algorithms to evaluate detection accuracy.
  • Identified three consensus biomarkers via feature importance analysis and evaluated their diagnostic reliability.
  • The integrated models achieved an AUC approaching 1.00 in distinguishing ovarian cancer from healthy samples.
  • Three specific mass-to-charge ratios were identified as consensus biomarkers for diagnostic use.
  • The study's model performance was supported by robust multi-dimensional feature contributions.

Abstract

Ovarian cancer poses a significant clinical challenge due to its asymptomatic onset and poor prognosis, highlighting the critical need for effective early detection strategies. This study developed a framework that integrates serum proteomic profiling with machine learning algorithms. Serum samples from 188 patients and 208 healthy controls were analysed via Matrix-Assisted Laser Desorption/Ionization Time - of - Flight (MALDI - TOF) mass spectrometry, revealing 43 differentially expressed peptides (17 upregulated, 26 downregulated). An ensemble pipeline incorporating eight machine learning algorithms showed favorable discriminatory ability in the study cohort, with the area under the receiver operating characteristic curve (AUC) of the integrated models approaching 1.00 for distinguishing ovarian cancer samples from healthy controls. We identified three consensus biomarkers mass-to-charge ratio (m/z) = 4211.41, 2881.50, 2662.15 through feature importance analysis, and their diagnostic reliability was supported by receiver operating characteristic curves optimization and decision curve analysis. Interpretability approaches integrating Shapley values and LIME indicated that the model’s high performance (AUC ≈ 1) was driven by robust multi-dimensional feature contributions. Cross-referencing with existing datasets suggested the PDE11A as a potential diagnostic biomarker. Collectively, this ensemble machine learning algorithms leveraged on serum proteomics shows promising potential for early detection of ovarian cancer, offering a strategy to mitigate the limitations of single-analyte biomarkers.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ding et al. (2026) studied this question.

synapsesocial.com/papers/6a002222c8f74e3340f9d11ahttps://doi.org/10.1038/s41598-026-51553-4
Ask AI
Helpful
Bookmark
Share
View Full Paper