We consider the scenarios that multi-agent cooperatively compete the task of collision-free formation cruise in a specific region. Considering mission complexity and real world constraints, we propose reinforcement learning-based solution for USVs and UAVs cooperative formation. Firstly, based on the curriculum learning, the complex formation control task is decomposed into two-stage training. In the first phase, the USV team is trained with PPO algorithm to track the target moving along a predetermined trajectory and avoid obstacles under the interference of waves. Subsequently, in the second stage, UAV team is trained with similar method. When UAV team is trained, the control strategies of USV team are fixed to the neural network obtained in the first stage. Combining with partial observable information, we design the reward function to make the USVs and UAVs learn policy and maintain a stable linear formation during movement. We validated the effectiveness of the proposed method with two-stage-simulation of USVs and UAVs in Unity environment. Compared to traditional control methods, the proposed method enables agents to learn effective strategies by interacting with the environment through a relatively simple training process without accurate mathematical model. This result simplifies the complexity of formation control and provides an easier solution for multi-agent formation control.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cooperative Formation Control of USVs and UAVs Based on Reinforcement Learning

  • Ting Wu,
  • Linqi Ye,
  • Xianglong Li,
  • Yan Peng

摘要

We consider the scenarios that multi-agent cooperatively compete the task of collision-free formation cruise in a specific region. Considering mission complexity and real world constraints, we propose reinforcement learning-based solution for USVs and UAVs cooperative formation. Firstly, based on the curriculum learning, the complex formation control task is decomposed into two-stage training. In the first phase, the USV team is trained with PPO algorithm to track the target moving along a predetermined trajectory and avoid obstacles under the interference of waves. Subsequently, in the second stage, UAV team is trained with similar method. When UAV team is trained, the control strategies of USV team are fixed to the neural network obtained in the first stage. Combining with partial observable information, we design the reward function to make the USVs and UAVs learn policy and maintain a stable linear formation during movement. We validated the effectiveness of the proposed method with two-stage-simulation of USVs and UAVs in Unity environment. Compared to traditional control methods, the proposed method enables agents to learn effective strategies by interacting with the environment through a relatively simple training process without accurate mathematical model. This result simplifies the complexity of formation control and provides an easier solution for multi-agent formation control.