The advancement of deep reinforcement learning (DRL) algorithms has created new avenues to solve multi-constraint problems. Typically, multi-constraint advice entails maximizing the intended goal while taking into account several restrictions. The guidance module is able to optimize the guidance command in real time by acting as an intelligent entity that is capable of gathering information and input from its surroundings. In this paper, we develop a multi-stage guiding architecture using the prediction and correction concept as our foundation. This study decouples the guidance law to realize the impact angle constraints in the first stage. In the second step, we fitted a predictor to estimate the flight’s time-to-go. Additionally, by adjusting the bias term of the proportional navigation guidance law, we set a corrector to control the impact time and angle. Subsequently, we select a nonlinear optimization function in order to restrict the field-of-view. The study offers a technique for producing the multi-constraint guidance command using DRL.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Deep Reinforcement Learning-Based Multi-constraint Guidance with Field-of-View Limitation

  • Yuhui Pu,
  • Yuru Bin,
  • Hui Wang,
  • Haorui Yang

摘要

The advancement of deep reinforcement learning (DRL) algorithms has created new avenues to solve multi-constraint problems. Typically, multi-constraint advice entails maximizing the intended goal while taking into account several restrictions. The guidance module is able to optimize the guidance command in real time by acting as an intelligent entity that is capable of gathering information and input from its surroundings. In this paper, we develop a multi-stage guiding architecture using the prediction and correction concept as our foundation. This study decouples the guidance law to realize the impact angle constraints in the first stage. In the second step, we fitted a predictor to estimate the flight’s time-to-go. Additionally, by adjusting the bias term of the proportional navigation guidance law, we set a corrector to control the impact time and angle. Subsequently, we select a nonlinear optimization function in order to restrict the field-of-view. The study offers a technique for producing the multi-constraint guidance command using DRL.