Output Feedback Q-Learning for a Non-Zero-Sum Game Problem in Building HVAC Control
摘要
Building heating, ventilating, and air conditioning (HVAC) systems have one of the largest energy footprint worldwide, which necessitates the design of intelligent control algorithms that improve the energy utilization while still providing thermal comfort. In this work, the authors formulate the HVAC equipment dynamics in the setting of a two-player non-zero-sum cooperative game, which enables two decision variables (mass flow rate and supply air temperature) to perform joint optimization of the control utilization and thermal setpoint tracking by simultaneously exchanging their policies. The HVAC zone serves as a game environment for these two decision variables that act as two players in a game. It is assumed that dynamic models of HVAC equipment are not available. Furthermore, neither the state nor any estimates of HVAC disturbance (heat gains, outside variations, etc.) are accessible, but only the measurement of the zone temperature is available for feedback. Under these constraints, the authors develop a new data-driven Q-learning scheme employing policy iteration and value iteration with a bias compensation mechanism that accounts for unmeasurable disturbances and circumvents the need of full-state measurement. The proposed algorithms are shown to converge to the optimal solution corresponding to the generalized algebraic Riccati equations (GAREs) in dynamic games.