This study presents a multi-agent approach to modeling deci-sion-making processes for cooperative-competitive game scenarios generated by long-tailed distributions. We introduce a novel simulation environment that aims to replicate real-world resource allocation challenges, particularly in healthcare settings. The environment incorporates dynamic changes in agents’ perceptions and utilizes heavy-tailed distributions to model unpredictable events. We adapt the Multi-Agent Deep Deterministic Policy Gradient algorithm to this environment, demonstrating its effectiveness in learning and strategy development. Our approach includes a role-based system and a two-component reward function that guides agents through the learning process. The results, evaluated using metrics such as cumulative reward, average episodic reward, and approximate Kullback-Leibler divergence, indicate successful learning and sufficent agent strategies. This research contributes to the field of multi-agent systems and resource management by providing a framework that can generalize learned policies to varying numbers of agents and resources, making it applicable to real-world scenarios with changing conditions.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cooperative-Competitive Decision-Making in Resource Management: A Reinforcement Learning Perspective

  • Artem Isakov,
  • Danil Peregorodiev,
  • Pavel Brunko,
  • Ivan Tomilov,
  • Natalia Gusarova,
  • Alexandra Vatian

摘要

This study presents a multi-agent approach to modeling deci-sion-making processes for cooperative-competitive game scenarios generated by long-tailed distributions. We introduce a novel simulation environment that aims to replicate real-world resource allocation challenges, particularly in healthcare settings. The environment incorporates dynamic changes in agents’ perceptions and utilizes heavy-tailed distributions to model unpredictable events. We adapt the Multi-Agent Deep Deterministic Policy Gradient algorithm to this environment, demonstrating its effectiveness in learning and strategy development. Our approach includes a role-based system and a two-component reward function that guides agents through the learning process. The results, evaluated using metrics such as cumulative reward, average episodic reward, and approximate Kullback-Leibler divergence, indicate successful learning and sufficent agent strategies. This research contributes to the field of multi-agent systems and resource management by providing a framework that can generalize learned policies to varying numbers of agents and resources, making it applicable to real-world scenarios with changing conditions.