Machine learning interatomic potentials (MLIPs) are typically developed for globally ordered homogeneous systems (GOHomS), which exhibit only minor local deviations from equilibrium configurations. Consequently, most existing MLIPs trained on GOHomS often perform inadequately when applied to locally ordered heterogeneous systems (LOHetS), e.g., substitutional alloying elements in multicomponent alloys. To describe doping alloy systems, we develop a fine-tuned MLIP based on the MACE foundation model, specifically tailored for Mo-based dilute alloys containing one or two out of 20 substitutional elements: Cr, Fe, Mn, Nb, Re, Ta, Ti, V, W, Y, Zr, Al, Zn, Cu, Ag, Au, Hg, Co, Ni, and Hf. The model is built on more than 7000 equilibrium and non-equilibrium structures derived from first-principles density functional theory (DFT) calculations. The optimized large-scale fine-tuned model attains state-of-the-art accuracy, with a mean absolute error (MAE) and root-mean-square error (RMSE) of 2.27 meV/atom and 3.79 meV/atom for energy predictions, and 13.83 meV/Å and 24.26 meV/Å for force predictions, respectively. Systematic evaluation under different data-splitting protocols shows that unknown element extrapolation remains challenging under strict dopant hold-out, whereas substantially improved accuracy can be achieved in partial-exposure transfer settings. The fine-tuned models reduce the MAE by approximately 7–10 times compared to models trained from scratch, and by 10–20 times relative to zero-shot foundation models. This performance gain remains consistent across varying dataset sizes (equilibrium vs. non-equilibrium structures) and model scales. Our work illustrates the efficacy of transfer learning from globally ordered homogeneous systems to locally ordered heterogeneous multicomponent alloy environments. However, direct transfer to entirely unknown elements remains challenging, especially when proxy embeddings are employed without fine-tuning. Thus, to achieve high accuracy without incurring additional cost, it is essential to include unknown elements in the training dataset while minimizing the number of configurations containing known elements. Moreover, the current findings are primarily validated for dilute Mo-based alloy systems. Extending this approach to more compositionally complex alloy spaces may necessitate additional data and further fine-tuning.
Fang et al. (Thu,) studied this question.