Lack of exploration capability is an important constraint on the performance of multi-agent deep reinforcement learning (MADRL) algorithms. Through evolutionary selection, humans can quickly explore unfamiliar environments and master survival skills by using the episodic memory in the brain. Inspired by biology, this paper proposes a multi-agent reinforcement learning exploration framework based on causal episodic memory and potential evolution (MACMPE) to improve the exploration capability of multi-agent systems in different types of task scenarios. First, we construct a causal episodic memory module that introduces causal learning computation and selects samples with high causal impact for the current training phase for policy updating to accelerate the policy learning process. Next, by combining the advantages of breadth search of evolutionary algorithms and deep search of deep learning, we build a potential evolution (PE) module to improve the ability of the multi-agent system to find the optimal solution in the environment. Then, we combine and embed the CM module and the PE module into the MADRL algorithm MAAC. Finally, we experiment with task scenarios including cooperative collection, command movement, and target navigation, and extend this framework to different MADRL algorithms. Experimental results show that the MADRL algorithms, combined with the framework proposed in this study, outperform the baseline algorithm regarding exploration capability and have better universality for the number of agents and scene categories.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MACMPE: Exploration Framework for Multi-agent Reinforcement Learning via Causal Episodic Memory and Potential Evolution

  • Liqiang Tian,
  • Peiliang Wu,
  • Qian Zhang,
  • Bingyi Mao,
  • Wenbai Chen

摘要

Lack of exploration capability is an important constraint on the performance of multi-agent deep reinforcement learning (MADRL) algorithms. Through evolutionary selection, humans can quickly explore unfamiliar environments and master survival skills by using the episodic memory in the brain. Inspired by biology, this paper proposes a multi-agent reinforcement learning exploration framework based on causal episodic memory and potential evolution (MACMPE) to improve the exploration capability of multi-agent systems in different types of task scenarios. First, we construct a causal episodic memory module that introduces causal learning computation and selects samples with high causal impact for the current training phase for policy updating to accelerate the policy learning process. Next, by combining the advantages of breadth search of evolutionary algorithms and deep search of deep learning, we build a potential evolution (PE) module to improve the ability of the multi-agent system to find the optimal solution in the environment. Then, we combine and embed the CM module and the PE module into the MADRL algorithm MAAC. Finally, we experiment with task scenarios including cooperative collection, command movement, and target navigation, and extend this framework to different MADRL algorithms. Experimental results show that the MADRL algorithms, combined with the framework proposed in this study, outperform the baseline algorithm regarding exploration capability and have better universality for the number of agents and scene categories.