<p>Path planning is a key task enabling Unmanned Aerial Vehicles (UAVs) to accomplish missions autonomously and safely. This paper introduces an innovative path planning approach based on the Deep-Reinforcement-Learning (DRL) concepts with an enhanced Dynamic Reward Function (DRF) using local environmental information. Although the capabilities of UAVs’ sensors are limited in collecting nearby environmental information, the path planning problem is formulated as a Partially Observable Markov Decision Process (POMDP). To overcome the challenges of partial environmental observability, a Q-neural network has been constructed with temporal memory, extracting vital information from the historical observations sequences. The state of the environment is encoded using a State Encoder (SE) which transforms the environmental information into a new input format for the designed Deep Q-learning Network (DQN) algorithm. Besides, an Adaptive Sampling (AS) improvement mechanism with two types of memory pools, for crucial and regular samples, is introduced to deal with a new variant of the DQN algorithm called Improved DQN (IDQN). Several comparisons with the most commonly state-of-the-art DRL algorithms, namely Deep Deterministic Policy Gradient (DDPG) and Proximal Policy Optimization (PPO), as well as with local path planning Dynamic Window Approach (DWA) and Artificial Potential Field (APF) techniques, are carried out to assess and highlight the competing performance of the proposed IDQN algorithm. An ablation study, examining the effect of each introduced improvement mechanism DQN + SE, DQN + AS and DQN + SE + AS, is performed. The IDQN’s sampling efficiency and convergence stability are evaluated through comparisons with the Uniform Experience Replay (UER) and Prioritized Experience Replay (PER) methods. The simulation results demonstrate the effectiveness and superiority of the proposed IDQN-based UAV path planning approach under the considered navigation scenarios with increased complexity and partial observability challenges.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improved deep Q-network-based UAV path planning in partially observable environments

  • Raja Jarray,
  • Soufiene Bouallègue

摘要

Path planning is a key task enabling Unmanned Aerial Vehicles (UAVs) to accomplish missions autonomously and safely. This paper introduces an innovative path planning approach based on the Deep-Reinforcement-Learning (DRL) concepts with an enhanced Dynamic Reward Function (DRF) using local environmental information. Although the capabilities of UAVs’ sensors are limited in collecting nearby environmental information, the path planning problem is formulated as a Partially Observable Markov Decision Process (POMDP). To overcome the challenges of partial environmental observability, a Q-neural network has been constructed with temporal memory, extracting vital information from the historical observations sequences. The state of the environment is encoded using a State Encoder (SE) which transforms the environmental information into a new input format for the designed Deep Q-learning Network (DQN) algorithm. Besides, an Adaptive Sampling (AS) improvement mechanism with two types of memory pools, for crucial and regular samples, is introduced to deal with a new variant of the DQN algorithm called Improved DQN (IDQN). Several comparisons with the most commonly state-of-the-art DRL algorithms, namely Deep Deterministic Policy Gradient (DDPG) and Proximal Policy Optimization (PPO), as well as with local path planning Dynamic Window Approach (DWA) and Artificial Potential Field (APF) techniques, are carried out to assess and highlight the competing performance of the proposed IDQN algorithm. An ablation study, examining the effect of each introduced improvement mechanism DQN + SE, DQN + AS and DQN + SE + AS, is performed. The IDQN’s sampling efficiency and convergence stability are evaluated through comparisons with the Uniform Experience Replay (UER) and Prioritized Experience Replay (PER) methods. The simulation results demonstrate the effectiveness and superiority of the proposed IDQN-based UAV path planning approach under the considered navigation scenarios with increased complexity and partial observability challenges.