Protein language models (pLMs) have been widely adopted for various protein and peptide related prediction tasks and demonstrated promising performance. However, short peptides are substantially underrepresented in commonly used high-capacity pLM training datasets, which may limit model effectiveness for peptide-specific tasks. This study aimed to deliver a lightweight and efficient peptide language model for accurate peptide representation and downstream prediction. To achieve this, we constructed PepBERT, a peptide-specific language model pretrained from scratch on four custom peptide datasets comprising between 1.97 and 19.2 million peptide sequences. Two versions, PepBERT-large (4.9 million parameters) and PepBERT-small (1.86 million parameters), were developed and comprehensively evaluated on nine peptide prediction tasks. Both variants achieved best performance on the UniParc dataset with 19.2 million peptide sequence, and PepBERT outperformed or matched the benchmark model ESM-2 (7.5 million parameters) on eight of nine tasks, while requiring substantially fewer computational resources. In terms of inference efficiency, PepBERT-large and PepBERT-small achieved more than 16 × and 25 × faster inference speeds, respectively, compared with ESM-2. These results demonstrate that PepBERT effectively captures the structural and functional properties of short peptides with high accuracy and low computational cost. By improving peptide representation and prediction, PepBERT can accelerate the discovery of food-derived bioactive peptides and support sustainable functional food innovation. All datasets, pretrained models, and codes are publicly available at https://github.com/dzjxzyd/PepBERT .
Du et al. (Sun,) studied this question.