Nonlinear optimal guidance problems with terminal constraints are often analytically intractable, and approximate solutions based on reinforcement learning or approximate dynamic programming generally fail to provide stability guarantees due to the presence of inherent approximation errors in neural networks. This paper proposes a robust incremental policy iteration algorithm for nonlinear optimal guidance problems. First, the incremental guidance problem is defined and an incremental policy iteration algorithm is designed to mitigate the initial instability of the classical policy iteration. Then, the boundary of the incremental guidance command is determined by integrating the Lyapunov stability theory into the policy improvement step, which ensures that the entire command is theoretically stable. Simulation results of a specific impact-angle-constrained guidance problem verify advantages of the developed method on efficiency, stability, and optimality.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Robust Incremental Learning of Approximate Dynamic Programming for Nonlinear Terminal Guidance

  • Han Wang,
  • Lin Cheng,
  • Shengping Gong

摘要

Nonlinear optimal guidance problems with terminal constraints are often analytically intractable, and approximate solutions based on reinforcement learning or approximate dynamic programming generally fail to provide stability guarantees due to the presence of inherent approximation errors in neural networks. This paper proposes a robust incremental policy iteration algorithm for nonlinear optimal guidance problems. First, the incremental guidance problem is defined and an incremental policy iteration algorithm is designed to mitigate the initial instability of the classical policy iteration. Then, the boundary of the incremental guidance command is determined by integrating the Lyapunov stability theory into the policy improvement step, which ensures that the entire command is theoretically stable. Simulation results of a specific impact-angle-constrained guidance problem verify advantages of the developed method on efficiency, stability, and optimality.