PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 17, 2026IEEE Transactions on Computational Biology and Bioinformatics0 citations

Robust CRISPR-Cas Protein Identification using Max-Margin Regularized Transformer Models

View Full Paper
BNBharani NammiSMSita Sirisha MadugulaVJVindi M. Jayasinghe-Arachchige

Key Points

  • This study aims to enhance the identification of CRISPR-Cas proteins using advanced deep-learning models.
  • Developed transformer encoder-based classification model and fine-tuned large protein language model.
  • Introduced max-margin regularization for improved model performance.
  • Evaluated classification accuracy on multiple protein comparison tasks.
  • FTPB model achieved accuracies of 99.06% for Cas9 vs. non-Cas and above 94% for other comparisons.
  • LSRMT model outperformed FTPB with accuracies exceeding 99% across tasks.
  • Max-Margin regularization improved model robustness and feature identification.

Abstract

The discovery of CRISPR-Cas system has significantly advanced genome editing, offering vast applications in medical treatments and life sciences research. Despite their immense potential, the existing CRISPR-Cas systems still face challenges concerning size, delivery efficiency, and cleavage specificity. Addressing these challenges requires a deeper understanding of CRISPR-Cas proteins to advance the design and discovery of novel Cas proteins. Here, we study CRISPR-Cas proteins extensively using deep-learning techniques to build classification models that can differentiate between Cas and non-Cas proteins, as well as identify subfamilies Cas9 and Cas12. We developed two types of deep learning models: 1) a transformer encoder-based classification model, trained from scratch; and 2) a large protein language model fine-tuned on ProtBert, pre-trained on more than 200 million proteins. To boost learning efficiency for the model trained from scratch, we introduced a novel margin-based loss function to maximize inter-class separability and intra-class compactness in protein sequence embedding latent space of a transformer encoder. Our results show that the Fine-Tuned ProtBert-based (FTPB) classification model achieved accuracies of 99.06%, 94.42%, 96.80%, 97.57% for Cas9 vs. non-Cas, Cas12 vs.non-Cas, Cas9 vs. Cas12, and multi-class classification of Cas9 vs. Cas12 vs. non-Cas proteins, respectively. The Latent Space Regularized Max-Margin Transformer (LSRMT) model achieved classification accuracies of 99.81%, 99.81%, 99.06%, and 99.27% for the same tasks, respectively. These results demonstrate the effectiveness of the proposed Max-Margin-based latent space regularization in enhancing model robustness and generalization capabilities. Remarkably, the LSRMT model, even when trained on a significantly smaller dataset, outperformed the fine-tuned state-of-the-art large protein model. The high classification accuracies achieved by the LSRMT model demonstrate its proficiency in identifying discriminative features of CAS proteins, marking a significant step towards advancing our understanding of CAS protein structures in future research endeavors.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Nammi et al. (2026) studied this question.

synapsesocial.com/papers/6a095a427880e6d24efe06afhttps://doi.org/10.1109/tcbbio.2026.3693528
Ask AI
Helpful
Bookmark
Share
View Full Paper