<p>This paper proposes a model-free online value iteration (VI) algorithm for solving stochastic linear quadratic control problems with ergodic cost functions, where the diffusion term in the dynamics equation is influenced by both the state and control variables. First, we propose an offline VI algorithm based on the idea of stochastic approximation. However, this algorithm requires prior knowledge of the system parameters, which are not always readily available. To overcome this limitation, we then develop a (partially) model-free online learning algorithm based on VI. This algorithm only requires a single system trajectory and does not need the initial control to be stabilizing. By exploiting the growth rate of Itô’s integrals to handle the stochastic term generated by multiplicative noise, we provide a rigorous proof of the algorithm’s convergence. Finally, a simulation example is presented to validate the convergence of the proposed algorithms.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An online value iteration method for stochastic linear quadratic control with multiplicative noise

  • Shumei Li,
  • Bing-Chang Wang,
  • Baoqiang Zhang

摘要

This paper proposes a model-free online value iteration (VI) algorithm for solving stochastic linear quadratic control problems with ergodic cost functions, where the diffusion term in the dynamics equation is influenced by both the state and control variables. First, we propose an offline VI algorithm based on the idea of stochastic approximation. However, this algorithm requires prior knowledge of the system parameters, which are not always readily available. To overcome this limitation, we then develop a (partially) model-free online learning algorithm based on VI. This algorithm only requires a single system trajectory and does not need the initial control to be stabilizing. By exploiting the growth rate of Itô’s integrals to handle the stochastic term generated by multiplicative noise, we provide a rigorous proof of the algorithm’s convergence. Finally, a simulation example is presented to validate the convergence of the proposed algorithms.