In the task offloading scenario of edge computing, user devices commonly suffer from resource constraints. It is difficult for traditional task offloading mechanisms to make efficient decisions when facing complex environments such as dynamically changing network resources and parallel scheduling of multiple tasks. Reinforcement learning methods face the problem of balance between exploration and exploitation in practical applications. To this end, we propose a multi-objective reinforcement learning mechanism based on SAE-PPO, which is an improvement of the traditional Proximal Policy Optimization (PPO) algorithm to achieve efficient task offloading. First, for the dynamics of network resources and the demand of multi-task scheduling, we design an Actor-Critic network model incorporating a multi-head self-attention (MHSA) mechanism. It can capture the complex association between task features and resource distribution, and thus optimize the matching strategy between tasks and resources. Second, to address the balance between exploration and exploitation, we construct an intelligent exploration strategy with entropy constraint by introducing an entropy regularization in the objective function. The strategy prevents premature convergence to a local optimum and adapts to different load and network conditions. Experimental results show that the proposed method outperforms the UCB1, SPEA/R, and NSGA-III algorithms in latency and energy consumption, achieving a 70% reduction in latency and a 46% reduction in energy consumption in complex MEC environments. This paper provides an effective solution for resource scheduling in complex edge environments.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-Objective Reinforcement Learning for Edge Task Offloading with Multi-Head Self-Attention and Entropy Constraint

  • Xiaoli Lu,
  • Gaizhi Guo,
  • Zongzuo Yu,
  • Pengjv Zhang

摘要

In the task offloading scenario of edge computing, user devices commonly suffer from resource constraints. It is difficult for traditional task offloading mechanisms to make efficient decisions when facing complex environments such as dynamically changing network resources and parallel scheduling of multiple tasks. Reinforcement learning methods face the problem of balance between exploration and exploitation in practical applications. To this end, we propose a multi-objective reinforcement learning mechanism based on SAE-PPO, which is an improvement of the traditional Proximal Policy Optimization (PPO) algorithm to achieve efficient task offloading. First, for the dynamics of network resources and the demand of multi-task scheduling, we design an Actor-Critic network model incorporating a multi-head self-attention (MHSA) mechanism. It can capture the complex association between task features and resource distribution, and thus optimize the matching strategy between tasks and resources. Second, to address the balance between exploration and exploitation, we construct an intelligent exploration strategy with entropy constraint by introducing an entropy regularization in the objective function. The strategy prevents premature convergence to a local optimum and adapts to different load and network conditions. Experimental results show that the proposed method outperforms the UCB1, SPEA/R, and NSGA-III algorithms in latency and energy consumption, achieving a 70% reduction in latency and a 46% reduction in energy consumption in complex MEC environments. This paper provides an effective solution for resource scheduling in complex edge environments.