<p>This paper presents a Q-learning method to solve stochastic linear quadratic Stackelberg games involving a leader and <i>N</i> followers where the system dynamics are unknown. The objective is to obtain the equilibrium policies by solving the coupled Hamilton-Jacobi-Bellman equations based on the leader-follower hierarchy. For each player, the Q-function containing unknown system parameters can be approximated by a critic neural network and the control policy can be approximated by an actor neural network. Then the tuning laws are given according to Bellman equations and gradient descent methods. An online model-free algorithm is developed and proven to converge almost surely for arbitrary control policies when the persistent excitation condition holds. Under some mild conditions, it is proven that the closed-loop system state and estimated weight errors are almost surely uniformly ultimately bounded. Finally, a numerical example is given to demonstrate the effectiveness of the proposed algorithm.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Stackelberg games for continuous-time stochastic linear quadratic systems via Q-learning

  • Ying Cao,
  • Bing-Chang Wang,
  • Bo Sun

摘要

This paper presents a Q-learning method to solve stochastic linear quadratic Stackelberg games involving a leader and N followers where the system dynamics are unknown. The objective is to obtain the equilibrium policies by solving the coupled Hamilton-Jacobi-Bellman equations based on the leader-follower hierarchy. For each player, the Q-function containing unknown system parameters can be approximated by a critic neural network and the control policy can be approximated by an actor neural network. Then the tuning laws are given according to Bellman equations and gradient descent methods. An online model-free algorithm is developed and proven to converge almost surely for arbitrary control policies when the persistent excitation condition holds. Under some mild conditions, it is proven that the closed-loop system state and estimated weight errors are almost surely uniformly ultimately bounded. Finally, a numerical example is given to demonstrate the effectiveness of the proposed algorithm.