<p>Autonomous exploration of confined environments remains a critical challenge across domains such as disaster response and subterranean inspection. Traditional methods such as frontier-based, sampling-based, and random walk strategies have offered viable solutions, yet face limitations in communication- and Global Positioning System (GPS)-denied environments. Recently, learning-based exploration has emerged as a promising approach to enhance the autonomy and adaptability of Micro-Aerial Vehicles (MAVs) in constrained environments. This systematic literature review investigates learning-based exploration strategies for single-MAV systems, with an emphasis on deep reinforcement learning (DRL). A comprehensive search was conducted across IEEE Xplore, ScienceDirect, and Web of Science, including preprints and early-access articles published up to June 30, 2025. The review excludes multi-robot systems, pure path planning, area coverage, obstacle avoidance, and planetary exploration, while prioritising search-and-rescue, inspection, and navigation-integrated exploration in GPS- or communication-denied confined environments. From a pool of 4,335 studies, 25 met the inclusion criteria. Our analysis reveals a strong preference for model-free DRL over model-based approaches and a growing interest in sim-to-real transfer techniques. Actor-critic methods, particularly Proximal Policy Optimisation (PPO), dominate due to their robustness and training stability. Despite progress, challenges persist in real-world deployment, reward design, and generalisation. This review synthesises trends, training paradigms, simulation environments, and reward structures used for MAV learning-based exploration, highlighting progress and persistent gaps. It lays a foundation for future work by identifying opportunities for hybrid approaches, real-world validation, and benchmark development to accelerate the safe and scalable deployment of learning-based MAV systems, specifically DRL ones in unknown environments. </p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Learning to Explore: A Systematic Review of Learning-Based Single MAV Exploration in Confined Environments

  • Makhosazana Eunice Moyo,
  • Turgay Celik

摘要

Autonomous exploration of confined environments remains a critical challenge across domains such as disaster response and subterranean inspection. Traditional methods such as frontier-based, sampling-based, and random walk strategies have offered viable solutions, yet face limitations in communication- and Global Positioning System (GPS)-denied environments. Recently, learning-based exploration has emerged as a promising approach to enhance the autonomy and adaptability of Micro-Aerial Vehicles (MAVs) in constrained environments. This systematic literature review investigates learning-based exploration strategies for single-MAV systems, with an emphasis on deep reinforcement learning (DRL). A comprehensive search was conducted across IEEE Xplore, ScienceDirect, and Web of Science, including preprints and early-access articles published up to June 30, 2025. The review excludes multi-robot systems, pure path planning, area coverage, obstacle avoidance, and planetary exploration, while prioritising search-and-rescue, inspection, and navigation-integrated exploration in GPS- or communication-denied confined environments. From a pool of 4,335 studies, 25 met the inclusion criteria. Our analysis reveals a strong preference for model-free DRL over model-based approaches and a growing interest in sim-to-real transfer techniques. Actor-critic methods, particularly Proximal Policy Optimisation (PPO), dominate due to their robustness and training stability. Despite progress, challenges persist in real-world deployment, reward design, and generalisation. This review synthesises trends, training paradigms, simulation environments, and reward structures used for MAV learning-based exploration, highlighting progress and persistent gaps. It lays a foundation for future work by identifying opportunities for hybrid approaches, real-world validation, and benchmark development to accelerate the safe and scalable deployment of learning-based MAV systems, specifically DRL ones in unknown environments.