Explainable reinforcement learning for enhancing personal thermal comfort and optimizing demand response in household multi-zone HVAC system
摘要
Heating, ventilation and air conditioning (HVAC) systems maintain personal thermal comfort (PTC) and serve as key demand response (DR) resources. Given the challenges of uncertain environments, varied user preferences, and legal requirements for transparent control strategies, this paper proposes an explainable reinforcement learning (XRL) solution for household multi-zone HVAC systems. This approach aims to optimize energy costs, ensure users’ PTC, and maintain the explainability of the DR strategy under uncertainty. Firstly, an XRL-based optimization framework is proposed. The framework utilizes XRL’s online learning capabilities to handle uncertainties and meet the PTC requirements of different zones while maintaining the explainability of the optimization strategy. Then, we propose an explainable proximal policy optimization (XPPO) algorithm as an instantiation of XRL for optimizing household multi-zone HVAC systems, using interpretable continuous control trees as actor networks of the XPPO. Moreover, we design state space, action space, reward, learning network, and learning algorithm of the XPPO in detail for the needs of HVAC operational optimization. Simulation results show that our method is explainable and thus can fulfill the requirements of the laws. At the same time, the proposed method consumes 22.4% less energy cost compared to scenarios without DR. Furthermore, the optimization of DR for a typical 4-zone household can be achieved within 15 min on an i7-12700 CPU.