<p>Learning effective strategies in sparse reward tasks is a fundamental challenge in reinforcement learning, which becomes especially difficult in multi-agent environments due to non-stationarity and exponentially growing joint state spaces induced by concurrent agent learning. While existing methods promote cooperation through experience sharing, learning from large collections of shared experiences proves inefficient in sparse reward settings where high-value states are rare, potentially exacerbating the curse of dimensionality in large-scale systems. This paper proposes effective multi-agent selective learning methods to boost sample-efficient training by learning from successful experiences. We adopt a retrogression-based selection method to identify successful agent trajectories from the team rewards, based on which some recall traces are generated and shared among agents to motivate effective exploration. Moreover, we selectively consider information from other agents to cope with the non-stationarity issue while enabling efficient training for large-scale agents. Considering scenarios with static and dynamic targets, we introduce vanilla MASL and MASL-DI algorithms, respectively. The former learns from recall traces by behavioral cloning, while the latter employs a dual imitation mechanism to improve exploration efficiency in dynamic target environments. Experimental results show that our method significantly improves sample efficiency in sparse reward tasks compared with baselines, especially in large-scale environments, achieving over 30% improvement in sample efficiency.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Learning from success: efficient selective learning methods for multi-agent sparse-reward tasks

  • Xinning Chen,
  • Xuan Liu,
  • Yanwen Ba,
  • Shigeng Zhang,
  • Bo Ding,
  • Kenli Li

摘要

Learning effective strategies in sparse reward tasks is a fundamental challenge in reinforcement learning, which becomes especially difficult in multi-agent environments due to non-stationarity and exponentially growing joint state spaces induced by concurrent agent learning. While existing methods promote cooperation through experience sharing, learning from large collections of shared experiences proves inefficient in sparse reward settings where high-value states are rare, potentially exacerbating the curse of dimensionality in large-scale systems. This paper proposes effective multi-agent selective learning methods to boost sample-efficient training by learning from successful experiences. We adopt a retrogression-based selection method to identify successful agent trajectories from the team rewards, based on which some recall traces are generated and shared among agents to motivate effective exploration. Moreover, we selectively consider information from other agents to cope with the non-stationarity issue while enabling efficient training for large-scale agents. Considering scenarios with static and dynamic targets, we introduce vanilla MASL and MASL-DI algorithms, respectively. The former learns from recall traces by behavioral cloning, while the latter employs a dual imitation mechanism to improve exploration efficiency in dynamic target environments. Experimental results show that our method significantly improves sample efficiency in sparse reward tasks compared with baselines, especially in large-scale environments, achieving over 30% improvement in sample efficiency.