This paper investigates the performance of a deep reinforcement learning (DRL) algorithm in robotic navigation, focusing on how environmental complexity and initial conditions affect learning efficiency. The study aims to identify which environmental factors contribute to faster or slower learning and how intrinsic exploration behavior influences navigation performance. A simulated environment was used to measure success rates, collision rates, time to completion, and the proportion of episodes in which the agent reached the goal or collided. The test scenarios varied the initial positions of the robot and goal (fixed or randomized) and the presence or absence of obstacles. The results indicate that the learning efficiency decreases significantly in obstacle-rich environments with randomized conditions, particularly when the robot and goal positions vary. In contrast, simpler cases with fixed starting conditions or a randomized robot position with a fixed goal led to faster convergence and higher success rates. This study provides a deeper understanding of DRL performance boundaries and highlights key factors influencing the speed and reliability of policy learning in robotic navigation tasks, emphasizing the role of intrinsic exploration driven by a stochastic policy.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating Deep Reinforcement Learning for Robotic Navigation

  • Jorge Gutiérrez,
  • Luis Enrique Sucar,
  • Rafael Murrieta-Cid,
  • Israel Becerra,
  • Gabriel Omar Flores-Aquino

摘要

This paper investigates the performance of a deep reinforcement learning (DRL) algorithm in robotic navigation, focusing on how environmental complexity and initial conditions affect learning efficiency. The study aims to identify which environmental factors contribute to faster or slower learning and how intrinsic exploration behavior influences navigation performance. A simulated environment was used to measure success rates, collision rates, time to completion, and the proportion of episodes in which the agent reached the goal or collided. The test scenarios varied the initial positions of the robot and goal (fixed or randomized) and the presence or absence of obstacles. The results indicate that the learning efficiency decreases significantly in obstacle-rich environments with randomized conditions, particularly when the robot and goal positions vary. In contrast, simpler cases with fixed starting conditions or a randomized robot position with a fixed goal led to faster convergence and higher success rates. This study provides a deeper understanding of DRL performance boundaries and highlights key factors influencing the speed and reliability of policy learning in robotic navigation tasks, emphasizing the role of intrinsic exploration driven by a stochastic policy.