A cooperative dynamic target search approach for multi-UAV systems utilizing the MAPPO algorithm
摘要
With the continuous development of unmanned aerial vehicle (UAV) technology, the advantages of employing multiple UAVs to collaboratively perform complex tasks have become increasingly evident and have attracted significant attention from the industry. Traditional deep reinforcement learning models face challenges such as difficulty in convergence and long training times, which hinder their direct application to multi-UAV cooperative systems. In response to these issues, this study proposes an improved multi-agent proximal policy optimization algorithm (AS-MAPPO). This algorithm enables agents to share local observation information through communication, incorporates an attention mechanism to highlight important information, and introduces an exploration reward function. These enhancements allow agents to learn superior strategies in complex environments, significantly improving the efficiency and success rate of multi-UAV cooperative target search tasks. Simulation results demonstrate that the proposed algorithm outperforms MAPPO and MADDPG in terms of task success rate and execution time.