<p>In the paradigm shift towards intelligent air combat, the autonomous decision-making capability of unmanned combat aerial vehicles (UCAVs) has become a pivotal factor in achieving air superiority. However, existing deep reinforcement learning (DRL) agents exhibit a significant disconnect between strategic planning and tactical execution. They often lack strategic foresight due to difficulties in processing long-term dependencies, while their tactical maneuvers frequently become simplistic and predictable as they converge to deterministic policies. These fundamental deficiencies severely limit their combat effectiveness in complex and dynamic engagements.To address this core challenge, this thesis introduces PPO-ITDS, an integrated maneuver decision-making framework. The framework synergistically combines the Informer architecture for long-horizon strategic reasoning, a novel tanh-based objective with a double exponential weighted moving average (DEWMA) update mechanism for diverse and stable tactical exploration, and the sharpness-aware minimization lookahead (SAML) optimizer to enhance generalization.Extensive experimental results demonstrate the superiority of our approach. PPO-ITDS achieves a 43% increase in Elo score over the baseline PPO and secures win rates exceeding 70% against other mainstream DRL algorithms. This study concludes that a holistic framework, which systematically bridges the gap between strategic planning and tactical execution, is crucial. It presents a robust paradigm for developing the next generation of high-performance autonomous air combat systems.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Adaptive exploration and temporal attention in reinforcement learning for autonomous air combat decision making

  • Xiang Wu,
  • Junzhe Jiang,
  • Zhihong Chen,
  • Shaojie Wu,
  • Chenghong Ye,
  • Xueyun Chen

摘要

In the paradigm shift towards intelligent air combat, the autonomous decision-making capability of unmanned combat aerial vehicles (UCAVs) has become a pivotal factor in achieving air superiority. However, existing deep reinforcement learning (DRL) agents exhibit a significant disconnect between strategic planning and tactical execution. They often lack strategic foresight due to difficulties in processing long-term dependencies, while their tactical maneuvers frequently become simplistic and predictable as they converge to deterministic policies. These fundamental deficiencies severely limit their combat effectiveness in complex and dynamic engagements.To address this core challenge, this thesis introduces PPO-ITDS, an integrated maneuver decision-making framework. The framework synergistically combines the Informer architecture for long-horizon strategic reasoning, a novel tanh-based objective with a double exponential weighted moving average (DEWMA) update mechanism for diverse and stable tactical exploration, and the sharpness-aware minimization lookahead (SAML) optimizer to enhance generalization.Extensive experimental results demonstrate the superiority of our approach. PPO-ITDS achieves a 43% increase in Elo score over the baseline PPO and secures win rates exceeding 70% against other mainstream DRL algorithms. This study concludes that a holistic framework, which systematically bridges the gap between strategic planning and tactical execution, is crucial. It presents a robust paradigm for developing the next generation of high-performance autonomous air combat systems.