Reward Shaping for Video Game Playing Agents Based on Human Motivations
摘要
This study examines the application of human gamer motivations to reinforcement learning agents tasked with playing Pokémon Red. Using Quantic Foundry's gamer motivation model, which categorizes gamers into nine types (Acrobat, Gardener, Slayer, Skirmisher, Gladiator, Ninja, Bounty Hunter, Architect, and Bard), we trained nine RL agents with reward functions weighted according to the motivational profiles of each gamer type. The agents were trained using Proximal Policy Optimization for approximately 100 million steps, with their performance measured through game progression metrics including badges earned, Pokémon levels, Pokédex completion, and trainers defeated. Results revealed distinct behavioral patterns among the agents, with the Gardener agent demonstrating superior performance across all progression metrics, while the Ninja and Bard agents consistently underperformed. These outcomes align with the expected behaviors of their human counterparts—Gardeners prioritize task completion, while Bards focus on exploration and fantasy over progression. Our findings suggest that incorporating human motivational profiles into reinforcement learning reward functions can produce agents that mimic the diversity of human approaches to gaming tasks.