A Value Iteration Algorithm for Stochastic Linear Quadratic Regulator
摘要
In this paper, we propose a novel value iteration algorithm for online adaptive optimal control of discrete-time stochastic linear quadratic regulator (LQR) problems. The algorithm iteratively solves the algebraic Riccati equation (ARE) using online information of states and inputs, without requiring the knowledge of the system dynamics. It does not require a discount factor because the issue of excitation noise bias is not present in our algorithm. Firstly, we review the optimal solution for the stochastic LQR problem. Secondly, we offer an offline model-based algorithm for solving ARE and prove its convergence. Thirdly, we present a data-driven online value iteration algorithm for solving ARE. Finally, we evaluate the proposed algorithm by an example.