Adaptive multi-player poker policy learning based on opponent style modeling
摘要
Multi-player games, such as multi-player poker, are a critical area of research and application in game theory. Currently, there is no theoretical optimal policy for multi-player game problems, including multi-player poker, and there is a lack of policy learning methods capable of adaptively making decisions against diverse opponents. To address this gap, this paper proposes an adaptive multi-player poker policy (AMP3) learning method based on opponent style modeling (OSM) for multi-player Texas Hold’Em poker games. First, we construct a style library and a gaming dataset for poker, designing style features based on traditional and statistical indicators specific to Texas Hold’Em. Next, we propose an OSM algorithm leveraging deep learning to predict style features from opponents’ historical data. Third, a novel reinforcement learning (RL) algorithm, based on the Actor-Critic framework, is introduced to learn adaptive policies by utilizing the opponents’ style features. To the best of our knowledge, this paper is the first to integrate OSM into RL to develop the AMP3 learning algorithm. Experimental results show that the proposed method exhibits strong adaptability in adjusting its policy style when facing players with varying strategies in six-player Texas Hold’Em poker.