<p>This paper considers the value iteration algorithms of stochastic zero-sum linear quadratic games with unkown dynamics. On-policy and off-policy learning algorithms are developed to solve the stochastic zero-sum games, where the system dynamics is not required. By analyzing the value function iterations, the convergence of the model-based algorithm is shown. The equivalence of several types of value iteration algorithms is established. The effectiveness of model-free algorithms is demonstrated by a numerical example.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

On-Policy and Off-Policy Value Iteration Algorithms for Stochastic Zero-Sum Dynamic Games

  • Liangyuan Guo,
  • Bing-Chang Wang,
  • Ji-Feng Zhang

摘要

This paper considers the value iteration algorithms of stochastic zero-sum linear quadratic games with unkown dynamics. On-policy and off-policy learning algorithms are developed to solve the stochastic zero-sum games, where the system dynamics is not required. By analyzing the value function iterations, the convergence of the model-based algorithm is shown. The equivalence of several types of value iteration algorithms is established. The effectiveness of model-free algorithms is demonstrated by a numerical example.