Peptide self-assembly underlies the formation of diverse supramolecular nanostructures with applications in biomaterials, drug delivery, and nanotechnology. We formulated a binary classification problem to predict self-assembly propensity using peptide sequences labeled as self-assembling (positive) or non-assembling (negative). Three feature representation strategies were evaluated: (1) amino-acid physicochemical descriptors, (2) natural language processing-inspired k-mer bag-of-words encodings, and (3) engineered sequence-derived descriptors. Each strategy was implemented using the Orange Data Mining platform with multiple supervised learning algorithms including logistic regression, support vector machines, AdaBoost, gradient boosting, random forests, and feedforward neural networks. Models were trained and assessed via stratified cross-validation on a balanced self-assembly dataset, and receiver operating characteristic (ROC) curves were generated from cross-validation predictions for the positive class (label 1, self-assembling peptides).2 The k-mer bag-of-words pipeline achieved the strongest performance with area under the ROC curve (AUC) values up to 0.998 and classification accuracy above 0.97.2 Across pipelines, ensemble methods and neural networks generally outperformed simpler models, although logistic regression remained competitive when paired with informative descriptors.2 These results support the hypothesis that sequence-only machine learning models can accurately predict peptide self-assembly and highlight the critical role of feature engineering in this task.
Daniel Tsen (Sat,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: