Obstacle avoidance in aerial robotics remains a critical challenge, particularly in environments with uncertain terrain and weather conditions. This study introduces a Constrained Q-Learning model that leverages spatio-temporal data from LiDAR and OctoMap to achieve Zero-Shot Execution (ZSE) for autonomous Unmanned Aerial Vehicle (UAV) navigation in unseen environments, eliminating the need for iterative training. Experimental evaluations are conducted using a high-fidelity simulator across three environments: random forests, clustered forests, and metropolitan areas, under varying obstacle densities and flight velocities. The proposed model demonstrates a 100% success rate, achieving average flight times of 45 s for slow velocities (below 2.5 m/s) and 34 s for fast velocities (above 2.5 m/s). Comparative analysis with 90 Human-to-Computer (HTC) flights (slow and fast velocities), conducted by three pilots under identical conditions, shows the proposed model reduces flight time by 33% (1.5 times faster) while enhancing path optimization. Additionally, the model matches the path selection efficiency of a standard Q-Learning approach without requiring iterative training, highlighting its robustness and scalability for autonomous UAV navigation in complex environments and GPS-denied locations.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Autonomous Navigation in Swarm of UAVs Using Spatio Temporal Data and Constrained-Reinforcement Learning

  • Abhudaya Shrivastava,
  • Christos Petridis,
  • Marijana Vacic,
  • Zoran Obradovic

摘要

Obstacle avoidance in aerial robotics remains a critical challenge, particularly in environments with uncertain terrain and weather conditions. This study introduces a Constrained Q-Learning model that leverages spatio-temporal data from LiDAR and OctoMap to achieve Zero-Shot Execution (ZSE) for autonomous Unmanned Aerial Vehicle (UAV) navigation in unseen environments, eliminating the need for iterative training. Experimental evaluations are conducted using a high-fidelity simulator across three environments: random forests, clustered forests, and metropolitan areas, under varying obstacle densities and flight velocities. The proposed model demonstrates a 100% success rate, achieving average flight times of 45 s for slow velocities (below 2.5 m/s) and 34 s for fast velocities (above 2.5 m/s). Comparative analysis with 90 Human-to-Computer (HTC) flights (slow and fast velocities), conducted by three pilots under identical conditions, shows the proposed model reduces flight time by 33% (1.5 times faster) while enhancing path optimization. Additionally, the model matches the path selection efficiency of a standard Q-Learning approach without requiring iterative training, highlighting its robustness and scalability for autonomous UAV navigation in complex environments and GPS-denied locations.