<p>In recent years, cooperative robot path planning has gained significant attention because of its wide range of applications, including load transportation and firefighting operations. However, existing methods based on classical control or hybrid strategies that combine control with deep reinforcement learning (DRL) often face challenges in generalizing and adapting to dynamic environments or unexpected situations. To address these issues, this work presents a fully DRL-based safe path planning approach for leader–follower robotic systems. In particular, we propose a modified Multi-Agent Twin-Delayed Deep Deterministic Policy Gradient (M-MATD3) algorithm, specifically designed to mitigate common issues such as overestimation bias and high variance observed in standard MATD3. Extensive simulations under a centralized training and decentralized execution (CTDE) framework show that the proposed M-MATD3 algorithm demonstrates strong performance, lowering collision rates to 8.15% in simple environments and 11.17% in complex ones, while significantly improving reward accumulation compared to existing techniques.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Deep reinforcement learning–based safe path planning for leader–follower robots

  • Ehsan Kazemi Tameh,
  • Mohammadreza Estarki,
  • Saeed Khodaygan

摘要

In recent years, cooperative robot path planning has gained significant attention because of its wide range of applications, including load transportation and firefighting operations. However, existing methods based on classical control or hybrid strategies that combine control with deep reinforcement learning (DRL) often face challenges in generalizing and adapting to dynamic environments or unexpected situations. To address these issues, this work presents a fully DRL-based safe path planning approach for leader–follower robotic systems. In particular, we propose a modified Multi-Agent Twin-Delayed Deep Deterministic Policy Gradient (M-MATD3) algorithm, specifically designed to mitigate common issues such as overestimation bias and high variance observed in standard MATD3. Extensive simulations under a centralized training and decentralized execution (CTDE) framework show that the proposed M-MATD3 algorithm demonstrates strong performance, lowering collision rates to 8.15% in simple environments and 11.17% in complex ones, while significantly improving reward accumulation compared to existing techniques.