Reference evapotranspiration (ET o ) is a critical climate parameter for drought management and the optimization of agricultural water use. This paper presents a machine learning-based framework for ET o prediction that was developed and tested at Santa Barbara Station, USA, and evaluated using data from four additional stations, namely Santa Monica, San Benito, San Diego II, and San Luis OW. Two feature selection approaches were employed: Particle swarm optimization (PSO)and a manually selected highly correlated variable (HCVS). These approaches identified the most effective input scenarios from 11 potential variables. The selected inputs were then used in machine learning models, including weighted instance handler wrapper (WIHW) combined with alternating model tree (AMTree) (WIHW-AMTree-PSO and WIHW-AMTree-HCVS), dual perturb and combine tree (WIHW-DPCTree-PSO and WIHW-DPCTree-HCVS), and random tree (WIHW-RANTree-PSO and WIHW-RANTree-HCVS). Analysis showed PSO reduced errors in a range of 4.06% to 30.77% compared with HCVS across all five stations. Among the models, WIHW-AMTree-PSO achieved the best performance, with root mean square error values of 0.213 mm/day at Santa Barbara, 0.235 mm/day at Santa Monica, 0.275 mm/day at San Benito, 0.239 mm/day at San Diego II, and 0.279 mm/day at San Luis OW. The corresponding percentage of Bias values were − 1.89%, −0.659%, 2.00%, 1.11%, and 2.82%, whereas Nash-Sutcliff efficiency values were 0.985, 0.975, 0.980, 0.972, and 0.966. Additionally, the Kling–Gupta efficiency values ranged from 0.966 to 0.982, and the legates–McCabe coefficient of efficiency values ranged from 0.867 to 0.900. Collectively, the results demonstrate that the proposed methodology offers high reliability for ET o prediction and holds considerable promise for broader applications in environmental, hydrological and climate-impact modeling. These findings highlight the value of advanced machine learning approaches for robust estimation of ET o , thereby supporting effective water resource management and agricultural planning, particularly in data-limited regions • Developed a machine learning framework for reference evapotranspiration prediction. • Models trained at one station and validated across four additional stations. • Two different techniques used to identify optimal input combinations. • Efficient inputs reduced prediction uncertainty by up to 31 %. • Models showed reliable performance and transferability to similar regions.
Khosravi et al. (Sun,) studied this question.