<p>Unmanned Aerial Vehicles need an online path planning capability to move in high-risk missions in unknown and complex environments to complete them safely. However, many algorithms reported in the literature may not return reliable trajectories to solve online problems in these scenarios. The Q-learning algorithm, a reinforcement learning technique, can generate trajectories in real-time and has demonstrated fast and reliable results. However, this technique has the disadvantage of defining the iteration number. If this value is poorly defined, it will take a long time or not return an optimal trajectory. Therefore, we propose a method to dynamically choose the number of iterations to obtain the best performance of Q-learning. The proposed method is compared to the Q-learning algorithm with a fixed number of iterations, A*, Rapid-Exploring Random Tree, Particle Swarm Optimization, Deep Q-learning, and Double Deep Q-Network. As a result, the proposed Q-learning algorithm demonstrates the efficacy and reliability of online path planning with a dynamic number of iterations to carry out online missions in unknown and complex environments. </p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dynamic Q-planning for online UAV path planning in unknown and complex environments

  • Lidia G. S. Rocha,
  • Kenny A. Q. Caldas,
  • Marco H. Terra,
  • Fabio Ramos,
  • Kelen C. Teixeira Vivaldini

摘要

Unmanned Aerial Vehicles need an online path planning capability to move in high-risk missions in unknown and complex environments to complete them safely. However, many algorithms reported in the literature may not return reliable trajectories to solve online problems in these scenarios. The Q-learning algorithm, a reinforcement learning technique, can generate trajectories in real-time and has demonstrated fast and reliable results. However, this technique has the disadvantage of defining the iteration number. If this value is poorly defined, it will take a long time or not return an optimal trajectory. Therefore, we propose a method to dynamically choose the number of iterations to obtain the best performance of Q-learning. The proposed method is compared to the Q-learning algorithm with a fixed number of iterations, A*, Rapid-Exploring Random Tree, Particle Swarm Optimization, Deep Q-learning, and Double Deep Q-Network. As a result, the proposed Q-learning algorithm demonstrates the efficacy and reliability of online path planning with a dynamic number of iterations to carry out online missions in unknown and complex environments.