Modified TD3 Reinforcement Learning-Based Path Following Control for an Autonomous Underwater Vehicle
摘要
This article aims to investigate the path following issue associated with the underactuated autonomous underwater vehicle (AUV), particularly in the presence of model uncertainties and disturbances from marine environments. A modified twin delayed deep deterministic policy gradient (TD3) algorithm is proposed to design a deep reinforcement learning controller. The prioritized experience replay mechanism is developed by calculating the sample priority and sampling probability. Double experience replay buffers using the successful and failed experience are constructed, and the adaptive proportionality coefficient is designed to adjust the composition structure of each mini-batch replay data. Besides, the long short-term memory layer is added to the deep network structure to handle the historical state of path following, thereby enhancing the training efficiency of the proposed modified TD3 algorithm. Simulation validations and comparative analyses demonstrate the significant efficacy and predominance of the proposed control method.