In military operational research, the interception task assignment problem stands as a pivotal decision optimization issue, primarily concerned with determining the optimal allocation scheme for intercepting the incoming targets. Traditional approaches to task assignment heavily relies on precise assessments of interception probability and threat levels. Conversely, recent years have seen the emergence of reinforcement learning based methodologies, which directly compute optimal allocation schemes through environmental state feedback. However, in most reinforcement learning based assignment algorithms, the neural network architecture constrains the agent’s adaptability in scenarios characterized by uncertain quantities of targets and interceptors. This paper introduces the Assigent algorithm, integrating the Transformer encoder module into the neural network architecture. The self-attention mechanism is utilized to extract vital information from state sequences involving unpredictable quantities of targets and interceptor launch platforms. Experimental results have shown that Assigent consistently outperforms both simple fully connected networks and conventional optimization methods in stochastic environments.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Self-attention Mechanism Based Reinforcement Learning for Task Assignment in Interception Scenarios

  • Bingchen Cai,
  • Naimin Zhang,
  • Han Yu

摘要

In military operational research, the interception task assignment problem stands as a pivotal decision optimization issue, primarily concerned with determining the optimal allocation scheme for intercepting the incoming targets. Traditional approaches to task assignment heavily relies on precise assessments of interception probability and threat levels. Conversely, recent years have seen the emergence of reinforcement learning based methodologies, which directly compute optimal allocation schemes through environmental state feedback. However, in most reinforcement learning based assignment algorithms, the neural network architecture constrains the agent’s adaptability in scenarios characterized by uncertain quantities of targets and interceptors. This paper introduces the Assigent algorithm, integrating the Transformer encoder module into the neural network architecture. The self-attention mechanism is utilized to extract vital information from state sequences involving unpredictable quantities of targets and interceptor launch platforms. Experimental results have shown that Assigent consistently outperforms both simple fully connected networks and conventional optimization methods in stochastic environments.