Wireless Resource Optimization for UAV Swarm Cooperative Sensing via Multi-agent Multi-task Deep Reinforcement Learning
摘要
Cooperative sensing through multiple unmanned aerial vehicles (UAV) swarms is a promising solution for many critical scenarios such as disaster monitoring, environmental surveillance, and military operations. Optimizing wireless resources, including spectrum and power allocation, is essential in UAV swarm systems to maximize their utilization. In this work, we investigate wireless resource optimization for a latency-sensitive multi-swarm coordination system based on deep reinforcement (DRL). A multi-objective joint subchannel and power allocation problem is formulated to minimize the expected age of information (AoI) and power consumption for each swarm, while maximizing the transmission probability of cooperative awareness messages (CAMs). The multi-agent deep deterministic policy gradient (MADDPG) framework is employed to simultaneously and effectively optimize multiple objectives, with each swarm leader (SL) acting as an independent agent capable of dynamically interacting with its environment to learn the optimal policy. Two distinct critic networks are designed to tackle the dual challenges of promoting cooperation among agents while also enhancing individual performance. A global critic is implemented to assess the system-wide expected reward, thereby encourage coordinated behaviors among agents. In parallel, local critics are tailored for each individual agent, focusing on evaluating agent-specific rewards. We also adopt a task-aware rewarding mechanism by decomposing the individual reward of each agent into multiple sub-reward functions based on the tasks each agent needs to accomplish, enabling the learning of task-specific value functions separately. Simulation results show that the proposed algorithm outperforms previous DRL frameworks in terms of AoI and CAM transmission.