In recent years, there has been a notable advancement in the field of Multi-Agent Reinforcement Learning (MARL). The Multi-Agent Transformer (MAT) represents the multi-agent decision-making problem as a sequence modeling problem and introduces the Transformer as a solution to the cooperative MARL problem. Mamba is a promising efficient State Space Model (SSM) that exhibits comparable performance to Transformers while having fewer parameters on long sequences. To enhance the efficacy of MARL in heterogeneous-agent and complex scenarios, we propose a more efficient and adaptable solution, the MA-Mamba method based on Mamba. In order to solve the problem of sparse temporal information and difficulty in modeling agent relationships in MARL, we propose to introduce a multi-channel mechanism to enhance the ability of MA-Mamba model to extract high-dimensional information. Moreover, we have implemented a hybrid module that integrates Mamba and attention mechanisms to facilitate the efficient encoding and decoding of information, thereby enabling the generation of more precise action sequences. Our method was tested on the StarCraftII Multi-Agent Challenge (SMAC), a widely utilized multi-Agent testing platform for MARL. The experimental results demonstrate that MA-Mamba achieves the best performance on all 13 tested maps, with four of them outperforming the best algorithm MAT.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MA-Mamba: Multi-Agent Reinforcement Learning with State Space Model

  • Yuan Lei,
  • Jian Xue,
  • Lin Zhao,
  • Chao Yue,
  • Ke Lu

摘要

In recent years, there has been a notable advancement in the field of Multi-Agent Reinforcement Learning (MARL). The Multi-Agent Transformer (MAT) represents the multi-agent decision-making problem as a sequence modeling problem and introduces the Transformer as a solution to the cooperative MARL problem. Mamba is a promising efficient State Space Model (SSM) that exhibits comparable performance to Transformers while having fewer parameters on long sequences. To enhance the efficacy of MARL in heterogeneous-agent and complex scenarios, we propose a more efficient and adaptable solution, the MA-Mamba method based on Mamba. In order to solve the problem of sparse temporal information and difficulty in modeling agent relationships in MARL, we propose to introduce a multi-channel mechanism to enhance the ability of MA-Mamba model to extract high-dimensional information. Moreover, we have implemented a hybrid module that integrates Mamba and attention mechanisms to facilitate the efficient encoding and decoding of information, thereby enabling the generation of more precise action sequences. Our method was tested on the StarCraftII Multi-Agent Challenge (SMAC), a widely utilized multi-Agent testing platform for MARL. The experimental results demonstrate that MA-Mamba achieves the best performance on all 13 tested maps, with four of them outperforming the best algorithm MAT.