To address the challenge of coordinating economic and environmental objectives for Multi-energy Virtual Power Plants (MEVPPs), particularly under ambitious decarbonization policies such as China’s “dual carbon” goals, this paper proposes a novel two-stage scheduling framework that integrates Deep Reinforcement Learning (DRL) with Model Predictive Control (MPC). The core innovations include the following: (1) high-fidelity physical models capturing wind turbulence correction, photovoltaic temperature-irradiation coupling, and state-of-charge-dependent energy storage efficiency, improving equipment dynamic characterization accuracy by 12.7% compared to conventional models; (2) an enhanced Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm incorporating priority experience replay and adaptive noise exploration, which accelerates convergence by 15.6%; (3) a pioneering coordination architecture of “Day-Ahead MADDPG—Real-Time MPC” that manages uncertainties through bidirectional feedback, where real-time deviations refine the long-term policy via experience replay. Simulation results using historical data from a North China industrial park demonstrate that the framework reduces operating costs by 13.3% and carbon emissions by 17.7% compared to particle swarm optimization, outperforms standard DDPG with 3.2% lower operating costs, 5.8% lower carbon emissions, and a 3.3% higher renewable utilization rate (88.6%), and achieves 55% renewable penetration with only 4.1% curtailment. These results validate the framework’s scalability for high-renewable penetration grids and its real-time feasibility, as confirmed by edge computing deployment with latency below 50 ms. This study offers a technically viable and scalable solution for the operation of low-carbon virtual power plants (VPPs), supporting the transition towards sustainable power systems.
Ni et al. (Wed,) studied this question.