A deep reinforcement learning framework for UAV navigation and selection of relay paths
摘要
The given paper proposes a deep reinforcement learning framework for unmanned aerial vehicles (UAVs) navigation to provide optimal communication for the users. It comprises gated recurrent unit enhanced with graph attention mechanism to make effective use of inter-UAV communication network for enhanced data retrieval and decisions. To further improve exploration and robustness, the proposed framework is trained using maximum-entropy reinforcement learning, enabling UAVs to learn stochastic policies. A heuristic reward function, solely based on local observations, is designed to optimize global performance in terms of coverage, fairness, and energy efficiency. Extensive simulations validate the effectiveness of the approach, showing superior scalability, adaptability, and communication efficiency over existing methods.