This study proposes a data-driven modelling alternative to conventional Density Functional Theory (DFT) calculations by employing three explainable ensemble learning algorithms (CatBoost, XGBoost, and Random Forest) to predict and interpret mechanical properties (bulk, shear, and Young's moduli) of ABX₃ perovskites based on elemental descriptors. All models demonstrated strong predictive performance, achieving R 2 values of ≥0.93 in both the training and testing phases. Among them, the Random Forest model consistently outperformed the others, yielding the highest R 2 scores, along with the lowest mean absolute error (MAE) and root mean squared error (RMSE), indicating superior generalization and robustness across all target properties. To enhance interpretability, Shapley Additive Explanations (SHAP) was used to quantify the contribution of individual features to model predictions. Feature importance analysis revealed that the melting points of the A- and B-site elements are among the most influential predictors of bulk modulus, likely due to their correlation with the material's resistance to thermal vibrations and mechanical deformation. Overall, this work demonstrates that explainable ensemble learning models, particularly Random Forest, can serve as scalable and cost-effective tools for accurately predicting and interpreting mechanical properties in perovskite materials. The integration of model interpretability with predictive insight offers a promising pathway for accelerating materials discovery and design in computational materials science. Correlation plots for the training phase for each of the ensemble learning methods. (a) Bulk; (b) shear; (c) Young moduli. • This study introduces an alternative approach to DFT by evaluating three ML models to predict and interpret the mechanical properties of perovskites. • CatBoost, XGBoost, and Random Forest models exhibited strong predictive performance. • The Random Forest model demonstrated superior performance comparatively. • The SHAP algorithm was applied to analyze the trained ensemble learning models for interpretability using a novel holistic method.
Akinpelu et al. (Sun,) studied this question.