In cooperative Multi-Agent Reinforcement Learning (MARL), agents must work together to accomplish a common goal. Intuitively, an agent will learn more efficiently if it considers other agents’ information contained in local observations and the temporal correlations of global states. Based on this motivation, in this paper, we propose Multi-Agent Temporal State Prediction and Sequence Recovery (T-SPSR), a novel MARL representation learning method. T-SPSR adopts different self-supervised learning signals for the policy and critic in MARL. Specifically, for the policy network, we introduce an auxiliary Prediction Model which treats the observations of all agents as a sequence and iteratively predicts agents’ future and previous observations. In this way, we expect the policy encoder can extract agent-level information and build awareness for other agents. For the critic network, we propose an auxiliary attention-based decoder to autoregressively recover the disordered global state sequence, expecting the critic encoder to capture temporal correlations between surrounding states. During training, the auxiliary model utilizes a self-supervised approach to co-optimize with the MARL’s loss. The experimental results on state-based and vision-based multi-agent benchmarks demonstrate that T-SPSR can greatly improve the performance and sample efficiency of MARL algorithm.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Temporal State Prediction and Sequence Recovery for Multi-agent Reinforcement Learning

  • Yong Wang,
  • Mingxiao Feng,
  • Haolin Song,
  • Wengang Zhou,
  • Houqiang Li

摘要

In cooperative Multi-Agent Reinforcement Learning (MARL), agents must work together to accomplish a common goal. Intuitively, an agent will learn more efficiently if it considers other agents’ information contained in local observations and the temporal correlations of global states. Based on this motivation, in this paper, we propose Multi-Agent Temporal State Prediction and Sequence Recovery (T-SPSR), a novel MARL representation learning method. T-SPSR adopts different self-supervised learning signals for the policy and critic in MARL. Specifically, for the policy network, we introduce an auxiliary Prediction Model which treats the observations of all agents as a sequence and iteratively predicts agents’ future and previous observations. In this way, we expect the policy encoder can extract agent-level information and build awareness for other agents. For the critic network, we propose an auxiliary attention-based decoder to autoregressively recover the disordered global state sequence, expecting the critic encoder to capture temporal correlations between surrounding states. During training, the auxiliary model utilizes a self-supervised approach to co-optimize with the MARL’s loss. The experimental results on state-based and vision-based multi-agent benchmarks demonstrate that T-SPSR can greatly improve the performance and sample efficiency of MARL algorithm.