ABSTRACT With the rapid development of wind power, photovoltaic systems and advanced energy storage technologies, integrating diverse renewable resources into existing power grids has become essential for improving operational efficiency and economic performance. Recognizing that wind, solar and storage units can be represented as configurable components within transmission network models, this paper formulates a transmission network expansion planning (TNEP) problem aimed at minimizing total investment costs under power demand constraints. To address this challenge, a novel approach combines reinforcement learning with Gaussian process regression (GPR) to approximate the Q ‐function in high‐dimensional, discrete action spaces. The GPR surrogate flexibly models nonlinear dependencies between expansion configurations and long‐term outcomes while quantifying uncertainty to guide focused exploration. This targeted learning strategy avoids exhaustive search and significantly improves efficiency, making it particularly suited to the combinatorial complexity of TNEP. Compared to linear regression‐based RL, which performs well only on small, smooth networks, the GPR‐based method achieves strong performance on both the Garver 6‐bus and IEEE 24‐bus systems. It consistently outperforms benchmark algorithms–including grey wolf optimizer, particle swarm optimization, genetic algorithm, and the gradient‐based BFGS method—in terms of convergence speed and solution quality, making it a practical tool for transmission expansion planning under high renewable penetration.
Shi et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: