The action selection strategy in multi-agent reinforcement learning faces the challenge of balancing exploration and exploitation. At the same time, homogeneous multi-agent systems, where all agents have the same structural characteristics, are widely present in the real world. Based on this characteristic, we propose a Probabilistic Action Selection Strategy(PASS) for homogeneous multi-agent reinforcement learning in balancing exploration and exploitation. This strategy represents the relationship between individual actions and global rewards as an action-reward function. By using the distribution formed by this function, the \(\varepsilon \) -greedy action selection strategy can make more valuable exploitation, reducing unnecessary exploration. We evaluated this in the StarCraft Multi-agent Challenge(SMAC) benchmark, and the experiments show that the improved probabilistic action selection strategy has a faster convergence speed and more stable performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Probabilistic \(\varepsilon \) -Greedy Strategy for Exploitation in Homogeneous Multi-agent Reinforcement Learning

  • Shanqi Li,
  • Xiangrui Meng,
  • Junqi Zhang,
  • Shaoqiu Zheng,
  • Yi Zuo,
  • Shixin Zheng,
  • Ying Tan

摘要

The action selection strategy in multi-agent reinforcement learning faces the challenge of balancing exploration and exploitation. At the same time, homogeneous multi-agent systems, where all agents have the same structural characteristics, are widely present in the real world. Based on this characteristic, we propose a Probabilistic Action Selection Strategy(PASS) for homogeneous multi-agent reinforcement learning in balancing exploration and exploitation. This strategy represents the relationship between individual actions and global rewards as an action-reward function. By using the distribution formed by this function, the \(\varepsilon \) -greedy action selection strategy can make more valuable exploitation, reducing unnecessary exploration. We evaluated this in the StarCraft Multi-agent Challenge(SMAC) benchmark, and the experiments show that the improved probabilistic action selection strategy has a faster convergence speed and more stable performance.