In response to the challenges encountered in complex marine settings, such as significant computational demands, imprecise control performance, constrained mobility, and challenges in counteracting environmental perturbations, we have implemented a strategy based on reinforcement learning for the control of target tracking by Unmanned Surface Vehicles (USV). By enhancing the network structure design and model training methods of the Proximal Policy Optimization (PPO) algorithm, we have improved its learning and representation capabilities in dealing with high-dimensional state spaces, thereby increasing the controller’s convergence speed, accuracy, and stability. Furthermore, an integral compensator was introduced to effectively eliminate the system’s steady-state error. The controller, based on the improved PPO algorithm, was trained in a simulation environment and subjected to comparative studies against controllers based on the Deep Deterministic Policy Gradient (DDPG), Twin Delayed DDPG (TD3), and traditional PPO algorithms. Simulation results demonstrate that our control strategy outperforms the other three methods, confirming the accuracy and effectiveness of our algorithm.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Target Tracking Control of Underactuated Unmanned Boats Using an Improved PPO Algorithm

  • Yuzhi Hao,
  • Qingling Wang,
  • Kaiyuan Shen

摘要

In response to the challenges encountered in complex marine settings, such as significant computational demands, imprecise control performance, constrained mobility, and challenges in counteracting environmental perturbations, we have implemented a strategy based on reinforcement learning for the control of target tracking by Unmanned Surface Vehicles (USV). By enhancing the network structure design and model training methods of the Proximal Policy Optimization (PPO) algorithm, we have improved its learning and representation capabilities in dealing with high-dimensional state spaces, thereby increasing the controller’s convergence speed, accuracy, and stability. Furthermore, an integral compensator was introduced to effectively eliminate the system’s steady-state error. The controller, based on the improved PPO algorithm, was trained in a simulation environment and subjected to comparative studies against controllers based on the Deep Deterministic Policy Gradient (DDPG), Twin Delayed DDPG (TD3), and traditional PPO algorithms. Simulation results demonstrate that our control strategy outperforms the other three methods, confirming the accuracy and effectiveness of our algorithm.