ABSTRACT Accurate prediction of polyethylene glycol (PEG) density is critical for its growing use as an eco‐friendly solvent in chemical separations and industrial processes, yet experimental measurements are time‐intensive and costly. This study addresses the challenge by developing a hybrid gradient boosting decision tree (GBDT) machine learning model to predict PEG density with high precision. The model's hyperparameters were optimized using four evolutionary algorithms including evolutionary strategies (ES), Bayesian probability improvement (BPI), batch Bayesian optimization (BBO), and self‐adaptive differential evolution (SADE) on a dataset of 293 points compiled from existing literature, with K ‐fold cross‐validation applied to prevent overfitting. Model performance was assessed via optimization runtime and metrics including R ‐squared ( R 2 ), mean squared error (MSE), and average absolute relative error percentage (AARE%). Key findings indicate that temperature has the strongest influence on PEG density (relevance factor: −0.87), followed by weaker correlations with pressure (0.39) and molecular weight (−0.17). The ES algorithm achieved the highest accuracy ( R 2 : 0.996, MSE: 1.149, AARE%: 0.078% on the test dataset), closely followed by BPI ( R 2 : 0.994, MSE: 1.663, AARE%: 0.092%), while SADE performed least effectively ( R 2 : 0.989, MSE: 1.939, AARE%: 0.091%) with the longest runtime (~2000 s). Sensitivity and SHAP analyses confirmed temperature's dominant role. These hybrid models offer a novel, computationally efficient alternative to experimental methods, advancing predictive modeling for PEG's physicochemical properties by integrating evolutionary optimization with GBDT, achieving superior accuracy compared to prior approaches.
Li et al. (2026) studied this question.