<p>In large-scale multiagent systems, the practical application of multiagent reinforcement learning (MARL) is hindered by the absence of robust reliability assurances. This gap has heightened the focus on strategy evaluation within the MARL framework, a domain that grapples with scalability issues in the joint strategy space. To address this concern, this paper introduces a novel two-stage graph-based strategy evaluation algorithm that significantly reduces the required sample capacity in the joint strategy space without compromising the evaluation quality. The proposed algorithm performs a hierarchical evaluation to compress sample capacity and employs a strategy-seeking model to seek a sink equilibrium (SE) joint strategy using the best responses. Moreover, a stopping condition is developed to achieve an approximately globally optimal SE strategy, accounting for the local optimal properties of the best-response-based algorithm. Case studies demonstrate that our algorithm achieves an approximately optimal SE joint strategy with superior sample efficiency compared with other approaches. The integration of MARL methods with the strategy evaluation algorithm proves to be an effective approach for establishing trustworthy MARL systems.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Graph-based strategy evaluation for large-scale multiagent reinforcement learning

  • Yiyun Sun,
  • Meiqin Liu,
  • Senlin Zhang,
  • Ronghao Zheng,
  • Shanling Dong

摘要

In large-scale multiagent systems, the practical application of multiagent reinforcement learning (MARL) is hindered by the absence of robust reliability assurances. This gap has heightened the focus on strategy evaluation within the MARL framework, a domain that grapples with scalability issues in the joint strategy space. To address this concern, this paper introduces a novel two-stage graph-based strategy evaluation algorithm that significantly reduces the required sample capacity in the joint strategy space without compromising the evaluation quality. The proposed algorithm performs a hierarchical evaluation to compress sample capacity and employs a strategy-seeking model to seek a sink equilibrium (SE) joint strategy using the best responses. Moreover, a stopping condition is developed to achieve an approximately globally optimal SE strategy, accounting for the local optimal properties of the best-response-based algorithm. Case studies demonstrate that our algorithm achieves an approximately optimal SE joint strategy with superior sample efficiency compared with other approaches. The integration of MARL methods with the strategy evaluation algorithm proves to be an effective approach for establishing trustworthy MARL systems.