• Samples with high information values are learned based on importance. • Conditional generate adversarial networks are combined with actor-critic networks. • Generative reinforcement learning method is proposed. • Dynamic optimal generation command allocation is solved. • Higher control performances with smaller frequency deviations are obtained. To improve the stability and suitability of high penetration of distributed energy resources, inspired by embedded reinforcement learning for generative artificial intelligence based on intelligent decision-making, this article presents low-carbon guidanced data-driven method, named as low-carbon guidanced generative reinforcement learning based on reinforcement learning and conditional generative adversarial networks. Firstly, the few-shot learning is trained to seek training samples with high information values; the conditional generative adversarial networks are trained to forecast the generation power of high penetration of distributed energy resources. Secondly, the conditional generative adversarial networks and few-shot learning connect with the actor-critic framework; the prediction mechanism of conditional generative adversarial networks is applied instead of the action selection mechanism of the actor-critic framework; the few-shot learning is utilized to learn important low carbon characteristics based on randomly generated low-carbon and high-performance control actions. Thirdly, the generative reinforcement learning can dynamically learn in complex multi-area novel power systems and generate optimal control commands considering low-carbon objectives for high penetration of distributed energy resources. The low-carbon guidanced generative reinforcement learning can reduce the average frequency deviation by 54.358% and reduce the average area control error by 33.767% with low carbon emissions. The integral squared error, integral absolute error, and integral time multiple absolute errors of average frequency deviations are at least smaller 31.635%, 33.377%, and 13.477% than other compared algorithms under benchmarking and real-life complexhigh penetration of distributed energy resources, respectively. Finally, the case studies show that the generative reinforcement learning have a faster convergence speed with greater control algorithms and low carbon emissions than other compared methods.
Li et al. (2026) studied this question.