Accurate prediction of concrete compressive strength is essential for structural safety and sustainable infrastructure development. Yet, conventional empirical approaches such as water-cement ratio laws and ACI/Eurocode formulations cannot fully capture the nonlinear interactions among modern concrete constituents, particularly supplementary cementitious materials and chemical admixtures. This paper presents ConcreteML, an open-source desktop application employing ensemble machine learning with engineering-informed compositional descriptors for interpretable and high-accuracy strength prediction. The framework integrates five machine learning algorithms (Random Forest, Gradient Boosting, XGBoost, CatBoost, and Neural Networks) trained on a multi-source dataset of 5186 concrete mixtures covering conventional, high-performance, and blended-cement concretes with curing ages from 1 to 365 days. The model incorporates 16 engineered descriptors, including water-cement ratio, total binder content, SCM percentage, and aggregate proportions, to transform raw mixture data into mechanistically meaningful inputs. The production XGBoost model achieved R² = 0.923, RMSE = 4.48 MPa, and MAE = 2.85 MPa on held-out test data, while 5-fold cross-validation (CV R² = 0.910 ± 0.008) confirmed its generalization. SHAP analysis provided transparent feature attribution, identifying age, total binder content, and water-cement ratio as the dominant predictors of strength development. ConcreteML includes a graphical user interface supporting single predictions, batch CSV/Excel processing, automated explainability visualizations, and standalone deployment without programming requirements.By combining predictive accuracy, explainability, and engineering usability, ConcreteML provides an accessible tool for quality control, mixture optimization, and data-driven concrete design.
Abbas et al. (Sun,) studied this question.