Collaborative Guidance Algorithm Based on Offline Pre-training and Online Reinforcement Learning
摘要
In response to the common assumption of small angle relationships in existing collaborative guidance laws and the neglect of high-order terms in the remaining time expansion, this paper proposes a guidance law structure based on a combination of traditional guidance laws and collaborative correction terms, and uses reinforcement learning methods to train the correction terms. This article also constructs a guided pre training algorithm based on offline reinforcement learning algorithms, combined with the dual delay deep deterministic policy gradient algorithm. Through methods such as delayed updates and critical comparison, fast and efficient learning and training iterations are carried out, effectively solving the problem of overestimation of actions and policies in the reinforcement learning process. The simulation results show that the reinforcement learning collaborative guidance law trained by the designed framework has obvious advantages of wider applicability and higher time collaboration accuracy.