Reinforcement-Learning-Based V2G Scheduling: Peak Load Mitigation and Financial Benefits
摘要
This study develops a centralized reinforcement learning framework, employing Q-learning, for efficient vehicle-to-grid (V2G) scheduling of large electric vehicle (EV) fleets. To address scalability challenges and inherent uncertainties in user behavior, the framework utilizes an aggregated state representation. This state captures the time-of-day, the distribution of the fleet’s state of charge (SOC) across discrete bins, and an estimated user adherence factor. The central agent learns a control policy based on this aggregated state to dynamically issue charging, discharging, or idle commands to EVs grouped within specific SOC bins. The primary objectives are to enhance power grid stability by minimizing the peak-to-average ratio (PAR) and to improve the economic viability of V2G participation for EV owners. Simulations conducted under a 40% EV penetration level (relative to a 300,000 vehicle base fleet) demonstrate the proposed method’s effectiveness: it reduced the grid PAR to 1.0683, compared to 1.0729 for a baseline uncontrolled charging scenario. Critically, the proposed method transformed the economic outcome, achieving positive average daily earnings of $1.17 per EV, in stark contrast to an average daily loss of $6.75 per EV under the baseline. These results validate the potential of the proposed intelligent, centralized control strategy using aggregated information to effectively manage large-scale V2G systems, enhance grid stability, and provide economic benefits under practical behavioral assumptions.