Dual-mode deep reinforcement learning for safety-oriented MASS collision avoidance
摘要
Safe collision avoidance for Maritime Autonomous Surface Ships (MASS) remains challenging because autonomous controllers must maintain predictable rule-guided behaviour while adapting to dense and uncertain traffic situations. This research introduces a dual-mode, safety-oriented deep reinforcement learning (DRL) framework that integrates model-based predictability with data-driven adaptability for MASS collision avoidance. Encounter scenarios are quantified in accordance with the International Regulations for Preventing Collisions at Sea (COLREGs), and a multi-objective reward function combines pairwise rule-guided reasoning with dynamic risk awareness. In the routine-navigation mode, Proximal Policy Optimisation (PPO) is employed to generate stable and COLREGs-guided trajectories, whilst the heightened-safety mode applies a tree-based safety filter that prunes unsafe actions and supports autonomous switching under elevated close-quarters risk. Two pruning optimisations–Reachable Envelope Pruning (REP) and the Terminal Safety Criterion (TSC)–jointly reduce node expansions by over 90%, thereby markedly improving computational efficiency. Simulation results demonstrate that the safety layer achieved no observed collisions in two-ship encounters and reduced collision rates by approximately 80–90% in congested multi-ship scenarios compared with the routine-navigation PPO baseline without the heightened-safety mode. Even in six-ship traffic, the agent maintains an average minimum passing distance exceeding 0.5 nautical miles. Latency analysis indicates reduced computational overhead under the tested simulation settings, suggesting the decision-level computational feasibility of the proposed framework.