<p>Wushu confrontation processes exhibit strong structural characteristics in terms of tactical organization, continuous action control, and knowledge dependence. Thus, there is a critical need to develop intelligent agents that can comprehensively understand tactical semantics and generate coherent, sequential actions. To solve this problem, this study proposes a Wushu tactical decision-making method based on Knowledge Graph-Guided Hierarchical Reinforcement Learning (KG-HRL), which combines tactical knowledge graph, Graph Attention Network (GAT) and hierarchical strategy structure into the same framework, so that tactical semantics, action correlation and strategy update can simultaneously act on the agent’s tactical decision-making process. In the simulated environment of Wushu confrontation, Knowledge Graph-Guided Hierarchical Reinforcement Learning (KG-HRL) is compared with baseline models such as Deep Q-Network (DQN), Proximal Policy Optimization (PPO), and Asynchronous Advantage Actor-Critic (A3C). The test consists of 1000 independent confrontations. The experiment adopts a strict win-loss determination protocol. If the scores of both sides are the same at the end of a regular round, a secondary determination is made based on the number of effective hits, the number of effective defenses, and the penalty for invalid actions. The test results only include two categories: win and loss. Draws are not included in the win rate. The strict win rate of KG-HRL is 72.3%. The average score is 68.5. The policy stability is 0.82. The action transition cost is 0.12. The tactical diversity index is 12. The control ability of offensive and defensive rhythm is 0.96. The interpretability of knowledge is 80.4%. Compared with the baseline model, KG-HRL demonstrates significant advantages under the same experimental protocol, with a significance level reaching <InlineEquation ID="IEq1"><EquationSource Format="TEX">\( p&lt;0.01\)</EquationSource></InlineEquation>. Compared with the baseline method, it shows advantages, <i>p</i> &lt; 0.01. It also demonstrates good adaptability under different styles of adversarial agents. The convergence time and calculation time of agent training are relatively balanced, and the training speed and reasoning efficiency are better. The training time is within 12.5&#xa0;h and the reasoning time is 3.8ms/ step. The ablation experiment shows that the performance of the agent decreases after the removal of each module, the knowledge graph and the attention structure of the graph have a high contribution to the state representation and tactical logic, and the hierarchical structure has a great influence on the generation of continuous actions. Different types of tactical relationships in knowledge graphs, such as offensive-defensive relationships and temporal constraints, contribute to decision-making performance to varying degrees. This study provides a systematic method for the modeling of Wushu intelligent agents based on knowledge structures, and offers certain reference significance for knowledge representation and strategy organization in complex action strategy tasks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Construction method of Wushu tactical decision-making agent driven by knowledge graph and reinforcement learning

  • Yufeng Ma

摘要

Wushu confrontation processes exhibit strong structural characteristics in terms of tactical organization, continuous action control, and knowledge dependence. Thus, there is a critical need to develop intelligent agents that can comprehensively understand tactical semantics and generate coherent, sequential actions. To solve this problem, this study proposes a Wushu tactical decision-making method based on Knowledge Graph-Guided Hierarchical Reinforcement Learning (KG-HRL), which combines tactical knowledge graph, Graph Attention Network (GAT) and hierarchical strategy structure into the same framework, so that tactical semantics, action correlation and strategy update can simultaneously act on the agent’s tactical decision-making process. In the simulated environment of Wushu confrontation, Knowledge Graph-Guided Hierarchical Reinforcement Learning (KG-HRL) is compared with baseline models such as Deep Q-Network (DQN), Proximal Policy Optimization (PPO), and Asynchronous Advantage Actor-Critic (A3C). The test consists of 1000 independent confrontations. The experiment adopts a strict win-loss determination protocol. If the scores of both sides are the same at the end of a regular round, a secondary determination is made based on the number of effective hits, the number of effective defenses, and the penalty for invalid actions. The test results only include two categories: win and loss. Draws are not included in the win rate. The strict win rate of KG-HRL is 72.3%. The average score is 68.5. The policy stability is 0.82. The action transition cost is 0.12. The tactical diversity index is 12. The control ability of offensive and defensive rhythm is 0.96. The interpretability of knowledge is 80.4%. Compared with the baseline model, KG-HRL demonstrates significant advantages under the same experimental protocol, with a significance level reaching \( p<0.01\). Compared with the baseline method, it shows advantages, p < 0.01. It also demonstrates good adaptability under different styles of adversarial agents. The convergence time and calculation time of agent training are relatively balanced, and the training speed and reasoning efficiency are better. The training time is within 12.5 h and the reasoning time is 3.8ms/ step. The ablation experiment shows that the performance of the agent decreases after the removal of each module, the knowledge graph and the attention structure of the graph have a high contribution to the state representation and tactical logic, and the hierarchical structure has a great influence on the generation of continuous actions. Different types of tactical relationships in knowledge graphs, such as offensive-defensive relationships and temporal constraints, contribute to decision-making performance to varying degrees. This study provides a systematic method for the modeling of Wushu intelligent agents based on knowledge structures, and offers certain reference significance for knowledge representation and strategy organization in complex action strategy tasks.