Adaptive exploration and temporal attention in reinforcement learning for autonomous air combat decision making
摘要
In the paradigm shift towards intelligent air combat, the autonomous decision-making capability of unmanned combat aerial vehicles (UCAVs) has become a pivotal factor in achieving air superiority. However, existing deep reinforcement learning (DRL) agents exhibit a significant disconnect between strategic planning and tactical execution. They often lack strategic foresight due to difficulties in processing long-term dependencies, while their tactical maneuvers frequently become simplistic and predictable as they converge to deterministic policies. These fundamental deficiencies severely limit their combat effectiveness in complex and dynamic engagements.To address this core challenge, this thesis introduces PPO-ITDS, an integrated maneuver decision-making framework. The framework synergistically combines the Informer architecture for long-horizon strategic reasoning, a novel tanh-based objective with a double exponential weighted moving average (DEWMA) update mechanism for diverse and stable tactical exploration, and the sharpness-aware minimization lookahead (SAML) optimizer to enhance generalization.Extensive experimental results demonstrate the superiority of our approach. PPO-ITDS achieves a 43% increase in Elo score over the baseline PPO and secures win rates exceeding 70% against other mainstream DRL algorithms. This study concludes that a holistic framework, which systematically bridges the gap between strategic planning and tactical execution, is crucial. It presents a robust paradigm for developing the next generation of high-performance autonomous air combat systems.