An experience knowledge driven learning algorithm based on deep reinforcement learning for long-distance UAV airdrop
摘要
It’s difficult for traditional solutions consisting of path planning and trajectory tracking algorithms to solve UAV maneuvering decision-making problem while executing long-distance UAV airdrop mission, because it’s necessary for obtaining an available path to prepare detailed information about mission area, such as digital elevation model, mountain, and other all possible threats. Therefore, it’s difficult for traditional methods to achieve UAV guidance in an interactive environment with an uncertain number of threats. An end-to-end algorithm is needed to perform long-distance UAV airdrop mission, where finding an optimal maneuvering decision-making policy becomes one of the key issues for improving the autonomy of UAV. In this paper, an experience knowledge driven learning algorithm based on deep reinforcement learning for long-distance UAV airdrop is proposed, which introduces expert experience to modify reward function for fully utilizing the potential value of experience transitions. Moreover, the training scheme of UAV maneuvering policy is designed based on Twin Delayed Deep Deterministic Policy Gradient with Prioritized Experience Replay (PER-TD3). Particularly, the training scheme incorporates Curriculum Learning (CL) to accelerate the convergence of policy. Compared with the traditional reward model without the inspiration of experience knowledge, our proposed algorithm could accelerate the convergence of policy and improve the performance of trained policy.