<p>Dynamic multi-robot task allocation (MRTA) requires real-time responsiveness and adaptability to rapidly changing conditions. Existing methods, primarily based on static data and centralized architectures, often fail in dynamic environments that require decentralized, context-aware decisions. To address these challenges, this paper proposes a novel graph reinforcement learning (GRL) architecture, named Spatial-Temporal Fusing Reinforcement Learning (STFRL), to address real-time distributed target allocation problems in search and rescue scenarios. The proposed policy network includes an encoder, which employs a Temporal-Spatial Fusing Encoder (TSFE) to extract input features and a decoder uses multi-head attention (MHA) to perform distributed allocation based on the encoder’s output and context. The policy network is trained with the REINFORCE algorithm. Experimental comparisons with state-of-the-art baselines demonstrate that STFRL achieves superior performance in path cost, inference speed, and scalability, highlighting its robustness and efficiency in complex, dynamic environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A graph reinforcement learning framework for real-time distributed multi-robot task allocation

  • Dian Zhang,
  • Peng Dong,
  • Pai Peng,
  • Yubo Dong

摘要

Dynamic multi-robot task allocation (MRTA) requires real-time responsiveness and adaptability to rapidly changing conditions. Existing methods, primarily based on static data and centralized architectures, often fail in dynamic environments that require decentralized, context-aware decisions. To address these challenges, this paper proposes a novel graph reinforcement learning (GRL) architecture, named Spatial-Temporal Fusing Reinforcement Learning (STFRL), to address real-time distributed target allocation problems in search and rescue scenarios. The proposed policy network includes an encoder, which employs a Temporal-Spatial Fusing Encoder (TSFE) to extract input features and a decoder uses multi-head attention (MHA) to perform distributed allocation based on the encoder’s output and context. The policy network is trained with the REINFORCE algorithm. Experimental comparisons with state-of-the-art baselines demonstrate that STFRL achieves superior performance in path cost, inference speed, and scalability, highlighting its robustness and efficiency in complex, dynamic environments.