Machine learning-based scoring functions have recently demonstrated great promise for large-scale virtual screening by processing 3D protein-ligand binding structures to predict binding affinities. However, their reliability is limited by the size and diversity of available training data sets. Data augmentation through modelling approaches, such as the introduction of the BindingNet data sets, has been shown to substantially improve model performance on pharmaceutically relevant benchmarks like the FEP benchmark. To advance data augmentation further, we propose NeuralBind, a large-scale training set constructed by modelling binding structures associated with high-quality ChEMBL activity records using Boltz-1x. Thanks to the de novo nature of co-folding models, NeuralBind eliminates the reliance on reference templates required in BindingNet construction and offers significantly greater diversity across the target space. With 177K entries, NeuralBind complements BindingNet in a non-overlapping way, enabling training scoring functions on nearly one million protein-ligand complexes. Additionally, this large scale of training data facilitates more advanced analyses, such as investigations into the trade-offs between data set volume and chemical breadth by comparing scoring functions trained on different NeuralBind subsets of equivalent size but varying chemical diversity. By providing a substantially expanded and diversified data set and informing future large-scale data curation strategies, this work paves the way for next-generation scoring functions with improved accuracy and generalizability.
Hsu et al. (Sun,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: