Enhanced perception and efficient exploration strategy for dynamic and uncertain UAV-assisted edge computing: AC-SAC
摘要
UAV-assisted Mobile Edge Computing (UAV-MEC) presents significant potential for efficient computing services, yet its dynamic resource management and task offloading face complex challenges, including high-dimensional continuous action spaces, environmental uncertainties, and decision coupling. Traditional Reinforcement Learning (RL) struggles with handling complex multi-entity states effectively, and under dynamic User Equipment (UE) relationships and challenging reward signals, it often exhibits low exploration efficiency and can get trapped in local optima. To address these issues, this paper proposes a novel AC-SAC (Attentive-Curious Soft Actor-Critic) algorithm. The algorithm is built upon the Soft Actor-Critic (SAC) framework. It first integrates a Multi-Head Self-Attention mechanism. This mechanism enables intelligent perception of complex multi-entity states, allowing the agent to adaptively weight state components and profoundly understand dynamic UE relationships. Consequently, it improves environmental evaluation and policy accuracy. Concurrently, the algorithm incorporates a Prediction Error-Based Curiosity Mechanism, which generates intrinsic rewards to effectively supplement external rewards, guiding efficient exploration, overcoming local optima, and significantly enhancing the policy’s exploration efficiency and robustness by deeply learning uncertain environmental dynamics. Experimental results demonstrate that the AC-SAC algorithm effectively tackles dynamic UAV-MEC environments. By jointly optimizing UAV trajectory, UE service selection, and task offloading, it reduces task processing delay, showcasing excellent performance and learning efficiency.