PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 2, 2026Briefings in Bioinformatics0 citationsOpen Access

Predicting protein–carbohydrate binding sites: a deep learning approach integrating protein language model embeddings and structural features

View Full Paper
MNMd Muhaiminul Islam NafiMRM Saifur Rahman

Key Points

  • The aim is to develop a deep learning model for predicting non-covalent protein-carbohydrate binding sites.
  • Built and tested an ensemble model, DeepCPBSite, combining three approaches: random undersampling, weighted oversampling, and class-weighted loss.
  • Utilized protein language model embeddings along with sequence and structural features for prediction.
  • Conducted SHAP analysis on structural features while categorizing proteins based on their organism information.
  • Achieved 78.7% balanced accuracy and 59.6% sensitivity on the TS53 set.
  • Outperformed the competitor, DeepGlycanSite, by 1.16% and 2.94% in accuracy and sensitivity, respectively.
  • Showed significant improvements in F1, MCC, and AUPR scores compared to state-of-the-art methods.

Abstract

Abstract Protein–carbohydrate interactions play an important role in many biological processes and functions, like inflammation, signal transduction, and cell adhesion. In our work, we will study non-covalent carbohydrate binding sites. In this paper, we aim to build a deep-learning model to predict non-covalent protein–carbohydrate binding sites. We were motivated by the fact that experimental approaches for predicting these sites are expensive. So, computational tools are necessary for identifying these interactions. We explored several sequence-based features as well as structural features. We also leveraged protein language model embeddings. We analyzed different architectures and selected the most suitable deep learning architecture for our finalized prediction model, DeepCPBSite. DeepCPBSite is an ensemble model that combines three separate models with three approaches (random undersampling, weighted oversampling, and class-weighted loss) built on the ResNet+FNN architecture. We made separate datasets from three sources: RCSB, UniProt, and CASP. We also compared the structural features extracted from the structures predicted by AlphaFold and ESMFold in the context of our prediction tasks. We employed three different feature selection techniques and finally did a SHAP (SHapley Additive exPlanations) analysis on the structural features after categorizing the proteins based on their organism information. DeepCPBSite achieved 78.7% balanced accuracy and 59.6% sensitivity on the TS53 set, outperforming the second-best competitor, DeepGlycanSite, by 1.16% and 2.94%, respectively. Additionally, its F1, MCC, and AUPR scores outperformed other state-of-the-art methods, with improvements ranging from 3.77%–47.6%, 3.84%–32.7%, and 8.18%–60.21%, respectively.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Nafi et al. (2026) studied this question.

synapsesocial.com/papers/6980fe48c1c9540dea810353https://doi.org/10.1093/bib/bbag008
Ask AI
Helpful
Bookmark
Share
View Full Paper