This study presents a machine learning framework for predicting and assessing membrane fouling risk in water treatment plants in Namibia. The approach integrates systematic data preprocessing with advanced predictive modelling techniques to estimate fouling propensity under varying physicochemical and operational conditions. Laboratory water quality data from Grunau, Opuwo, and Eenhana treatment plants were compiled and analysed using Decision Tree, Random Forest, Gradient Boosting, k Nearest Neighbours, and Support Vector Regression models. A continuous fouling severity index was developed, predicted, and analysed using time-series techniques to examine temporal patterns and site-specific dynamics. Model performance was evaluated using MAE, RMSE, R-squared, NSE, KGE, and PBIAS, supported by error distribution analysis and statistical testing. The results show that tree-based models consistently outperform other approaches, with Gradient Boosting achieving the best overall performance, with an R-squared of 0.962, an MAE of 0.793, and strong agreement across all evaluation metrics. Error distribution analysis confirmed that these models provide more stable and consistent predictions, while k-nearest neighbours and Support Vector Regression exhibit higher variability and reduced reliability. The findings indicate that early-stage fouling risk is primarily influenced by Total Dissolved Solids and pH, while breakthrough fouling events are associated with transient increases in iron and manganese concentrations. Time series analysis further revealed site-specific fouling cycles and temporal anomalies across treatment stages. The study demonstrates that tree-based machine learning models provide a reliable and practical tool for predicting fouling risk, supporting improved operational decision-making and proactive management in water treatment systems. • Tree-based models consistently outperformed other approaches, showing strong capability in capturing nonlinear and threshold-driven fouling behaviour. • Gradient Boosting achieved the best overall performance across all evaluation metrics, showing high accuracy, strong agreement, and stable error distribution. • Fouling dynamics are primarily controlled by Total Dissolved Solids and pH at early stages, while transient iron and manganese spikes drive breakthrough fouling events. • Error distribution and statistical tests confirmed significant differences in model performance, with tree-based models providing more reliable and consistent predictions. • Time series analysis revealed site-specific fouling patterns, including reduced scaling trends in Grunau, recurring metal-driven peaks in Opuwo, and consistently low fouling in Eenhana.
Ajayi et al. (Fri,) studied this question.