• A general machine learning strategy for imbalanced small-sample materials data. • B-SMOTE data augmentation reveals an optimal imbalance ratio. • Physically informed feature screening enhances interpretability and prediction. • Glass-forming ability predicted across wide amorphous alloy composition spaces. • Two new bulk metallic glasses designed and experimentally validated. In advanced materials systems, sample scarcity and class imbalance severely limit the generalization performance of data-driven models. To address this challenge, a strategy that combines data augmentation with dimensionality reduction of physicochemical features was proposed using amorphous alloys as a representative case. This approach enables high-precision prediction of glass-forming ability (GFA) in a high-dimensional compositional space and validates the feasibility of new alloy design. Specifically, we use the Borderline Synthetic Minority Oversampling Technique (B-SMOTE) to balance the dataset, while screening physicochemical factors to identify key alloy factors. A “key alloy factor–performance” model is then developed to predict GFA across the compositional space. The results demonstrate that a 2:1 majority-to-minority class ratio in the augmented data yields optimal model generalization. Further screening identifies ten critical alloy factors, and a support vector machine (SVM) model built on these factors achieves 89.09% accuracy on the test set, significantly outperforming existing conventional experience-driven or empirical-rule-based approaches. Based on this model, two bulk metallic glasses (Zr 38 Cu 28 Ag 7 Al 7 Be 20 and Zr 46 . 75 Cu 25 . 5 Al 12 . 75 Ni 11 Sc 4 ) are designed and successfully synthesized; their experimental characterization agrees well with the predictions. This study offers a new approach for addressing imbalanced small-sample problems in materials science and improving the reliability of predictions across broad composition spaces.
Zhu et al. (Sun,) studied this question.