Probabilistic \(\varepsilon \) -Greedy Strategy for Exploitation in Homogeneous Multi-agent Reinforcement Learning
摘要
The action selection strategy in multi-agent reinforcement learning faces the challenge of balancing exploration and exploitation. At the same time, homogeneous multi-agent systems, where all agents have the same structural characteristics, are widely present in the real world. Based on this characteristic, we propose a Probabilistic Action Selection Strategy(PASS) for homogeneous multi-agent reinforcement learning in balancing exploration and exploitation. This strategy represents the relationship between individual actions and global rewards as an action-reward function. By using the distribution formed by this function, the \(\varepsilon \) -greedy action selection strategy can make more valuable exploitation, reducing unnecessary exploration. We evaluated this in the StarCraft Multi-agent Challenge(SMAC) benchmark, and the experiments show that the improved probabilistic action selection strategy has a faster convergence speed and more stable performance.