<p>Deep reinforcement learning (DRL) is an effective approach to solving missile guidance problems. However, the broad exploration space and complex environment in the missile guidance process can make it difficult to improve its behavior in the early stage and remain stable in hard scenarios. For this reason, we propose a method named curriculum and imitation-based deep reinforcement learning (CIRL) for missile guidance. CIRL can provide action for the missile controller to hit a randomly maneuvering target with a great intersection angle and overcome detection delay and noise. To bootstrap and stabilize the agent’s training process, CIRL introduces double-adjust imitation learning (DAIL) to help the agent handle both good and bad exploration trajectories. Rule imitation learning is the first part of DAIL, enabling the agent to adapt to traditional guidance laws, avoid continued deterioration, and establish a basic policy in the early iterations. Gaussian self-imitative learning (GSIL) is introduced as the second part, which will help agents attach more importance to well-performed actions. We also apply curriculum learning to reduce the negative effect of imitation learning further promoting the agent’s exploration and enhancing robustness against various factors, including the target’s mobility, delay, and noise. Simulation results validate that CIRL outperforms traditional methods and state-of-the-art DRL-based guidance algorithms with a higher true-hit rate, which reaches 22.7%, in fewer iterations. The out-of-distribution (OOD) experiment is conducted to evaluate the robustness and susceptibility to overfitting, with results compared against PNG.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Rapid Bootstrapping of Deep Reinforcement Learning with Curriculum and Imitation Strategies for Missile Guidance

  • Runjian Xie,
  • Xing Jin,
  • Qinglei Zhao,
  • Yang Zhang,
  • Zhen Wang

摘要

Deep reinforcement learning (DRL) is an effective approach to solving missile guidance problems. However, the broad exploration space and complex environment in the missile guidance process can make it difficult to improve its behavior in the early stage and remain stable in hard scenarios. For this reason, we propose a method named curriculum and imitation-based deep reinforcement learning (CIRL) for missile guidance. CIRL can provide action for the missile controller to hit a randomly maneuvering target with a great intersection angle and overcome detection delay and noise. To bootstrap and stabilize the agent’s training process, CIRL introduces double-adjust imitation learning (DAIL) to help the agent handle both good and bad exploration trajectories. Rule imitation learning is the first part of DAIL, enabling the agent to adapt to traditional guidance laws, avoid continued deterioration, and establish a basic policy in the early iterations. Gaussian self-imitative learning (GSIL) is introduced as the second part, which will help agents attach more importance to well-performed actions. We also apply curriculum learning to reduce the negative effect of imitation learning further promoting the agent’s exploration and enhancing robustness against various factors, including the target’s mobility, delay, and noise. Simulation results validate that CIRL outperforms traditional methods and state-of-the-art DRL-based guidance algorithms with a higher true-hit rate, which reaches 22.7%, in fewer iterations. The out-of-distribution (OOD) experiment is conducted to evaluate the robustness and susceptibility to overfitting, with results compared against PNG.