APPLICATION OF REINFORCEMENT LEARNING IN A SHARED RESOURCE ENVIRONMENT REINFORCEMENT
摘要
Due to the growth of computing power, distributed systems are now widely used, combining many subsystems that operate freely in the environment and actively interact with each other. In this regard, the solution to the optimization problem of maximizing distributed resources is especially relevant. At the same time, it is necessary to ensure access to the current resources of all subsystems included in the system in dynamic conditions of mutual competition within the collective of subsystems as a whole. This problem is formulated using game theory as a placement game, in which players, knowing the contents of the payoff matrix and choosing strategies with different amounts of resources, compete for a resource that is relevant to all players within the team. At the same time, players strive to maximize its amount for themselves. The theory of collective behavior of automata allows us to successfully solve this problem. In this case, the machines are uniform elements of the system, forming a homogeneous collective. Optimal behavior in the collective is ensured by each machine using reinforcement learning with the help of rewards and penalties received from the environment, providing access to the distributed resource. In this paper, a comparison of the results of goal-directed behavior of groups of different types of machines is carried out. Various strategies of behavior of teams, depending on the inertial qualities of the machines that make up the team, are revealed. It is shown that the types of machines under study successfully solve the set task of maximizing distributed resources under conditions of competition within the team for a resource that is relevant for the entire team of machines. Computational experiments have shown that with the growth of memory, the machines, in terms of the sum of rewards and penalties from the environment, approach the actions of players who are fully informed about the values of the elements of the game’s payoff matrix.