PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 28, 2024BMC Bioinformatics29 citationsOpen Access

Protein embedding based alignment

View Full Paper
BIBenjamin Giovanni IovinoYYYuzhen Ye

Key Points

Key points are not available for this paper at this time.

Abstract

Abstract Purpose Despite the many progresses with alignment algorithms, aligning divergent protein sequences with less than 20–35% pairwise identity (so called "twilight zone") remains a difficult problem. Many alignment algorithms have been using substitution matrices since their creation in the 1970’s to generate alignments, however, these matrices do not work well to score alignments within the twilight zone. We developed Protein Embedding based Alignments, or PEbA, to better align sequences with low pairwise identity. Similar to the traditional Smith-Waterman algorithm, PEbA uses a dynamic programming algorithm but the matching score of amino acids is based on the similarity of their embeddings from a protein language model. Methods We tested PEbA on over twelve thousand benchmark pairwise alignments from BAliBASE, each one extracted from one of their multiple sequence alignments. Five different BAliBASE references were used, each with different sequence identities, motifs, and lengths, allowing PEbA to showcase how well it aligns under different circumstances. Results PEbA greatly outperformed BLOSUM substitution matrix-based pairwise alignments, achieving different levels of improvements of the alignment quality for pairs of sequences with different levels of similarity (over four times as well for pairs of sequences with <10% identity). We also compared PEbA with embeddings generated by different protein language models (ProtT5 and ESM-2) and found that ProtT5-XL-U50 produced the most useful embeddings for aligning protein sequences. PEbA also outperformed DEDAL and vcMSA, two recently developed protein language model embedding-based alignment methods. Conclusion Our results suggested that general purpose protein language models provide useful contextual information for generating more accurate protein alignments than typically used methods.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Iovino et al. (2024) studied this question.

synapsesocial.com/papers/68e771ffb6db6435876e6902https://doi.org/10.1186/s12859-024-05699-5
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1The language of proteins: NLP, machine learning & protein sequences2021 · 392 citations
  2. 2BAliBASE 3.0: Latest developments of the multiple sequence alignment benchmark2005 · 421 citations
  3. 3Deep embedding and alignment of protein sequences2022 · 62 citations
  4. 4Leveraging protein language models for accurate multiple sequence alignments2023 · 34 citations
  5. 5Evolutionary-scale prediction of atomic-level protein structure with a language model2023 · 5,468 citations