Abstract <p>In generating (sub-)optimal strategies for perfect information games, the dominant paradigm is reinforcement learning using neural networks to estimate actions (Q-values). The initially applied approach using a tree search with Monte Carlo evaluations was abandoned due to its lack of generalization ability. By applying a probabilistic-combinatorial formal learning method together with the Monte Carlo method we will show how generalized rules can be generated that form the desired strategy through similarities in states where the same action was applied leading to high reward.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Generating Strategies for Games with Complete Information Using Action Similarity

  • D. V. Vinogradov

摘要

Abstract

In generating (sub-)optimal strategies for perfect information games, the dominant paradigm is reinforcement learning using neural networks to estimate actions (Q-values). The initially applied approach using a tree search with Monte Carlo evaluations was abandoned due to its lack of generalization ability. By applying a probabilistic-combinatorial formal learning method together with the Monte Carlo method we will show how generalized rules can be generated that form the desired strategy through similarities in states where the same action was applied leading to high reward.