There have been many remarkable advances in the cooperative Multi-Agent Reinforcement Learning (MARL) area in recent years. Nevertheless, progress in model-based methods in MARL is limited due to the complexity of the world model compared to single-agent RL. The existing research, Model-based Value Decomposition (MBVD), is the only model-based method that combines the global state and imagination rollouts, achieving higher sample efficiency by providing foresight in centralized training. Yet, there is still room for improvement in the quality of these imaginations. In this research, we present a novel MARL model, called Enhanced Imagination and Feature Balancing (EIFB), which improves upon MBVD by introducing a new training objective of imagination module and a feature balancing network. We demonstrate the effectiveness of EIFB by integrating it with two MARL methods and evaluating it in the Predator-Prey and StarCraft Multi-Agent Challenge environments. Our experimental results indicate that our method achieves superior performance compared to MBVD, particularly with longer imagination horizons supported by robust feature integration.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhanced Imagination and Feature Balancing for Cooperative Multi-agent Reinforcement Learning

  • Fanchao Xu,
  • Tomoyuki Kaneko

摘要

There have been many remarkable advances in the cooperative Multi-Agent Reinforcement Learning (MARL) area in recent years. Nevertheless, progress in model-based methods in MARL is limited due to the complexity of the world model compared to single-agent RL. The existing research, Model-based Value Decomposition (MBVD), is the only model-based method that combines the global state and imagination rollouts, achieving higher sample efficiency by providing foresight in centralized training. Yet, there is still room for improvement in the quality of these imaginations. In this research, we present a novel MARL model, called Enhanced Imagination and Feature Balancing (EIFB), which improves upon MBVD by introducing a new training objective of imagination module and a feature balancing network. We demonstrate the effectiveness of EIFB by integrating it with two MARL methods and evaluating it in the Predator-Prey and StarCraft Multi-Agent Challenge environments. Our experimental results indicate that our method achieves superior performance compared to MBVD, particularly with longer imagination horizons supported by robust feature integration.