Dynamic Dual-Strategy Update of Offline-to-Online Reinforcement Learning for Optimal Energy System Scheduling
摘要
In modern power systems, efficient energy scheduling is essential for ensuring system stability and optimizing resource allocation. At present, Deep Reinforcement Learning (DRL) algorithms perform well in continuous action spaces, but they face challenges such as high online learning costs and poor offline learning performance. To address these challenges, this paper proposes a Dynamic Dual-Strategy Update (DDSU) DRL that combines Twin Delayed Deep Deterministic Policy Gradient (TD3) and Generative Adversarial Imitation Learning (GAIL) technologies. By dynamically updating the agent’s strategy, DDSU enables parallel training of offline and online reinforcement learning, thereby enhancing the efficiency and accuracy of energy system scheduling. Experimental results show that DDSU significantly improves the convergence speed and final performance in the energy system scheduling task compared to the TD3 algorithm, in terms of reward acquisition, cost control, and system stability. These results demonstrate the potential of DDSU in enhancing the efficiency and accuracy of power system scheduling.