Localized Affinity-Based Reinforcement Learning for Interpretable State-Specific Decision-Making
摘要
Designing a reward function that elicits the desired behavior poses a significant challenge in the field of reinforcement learning (RL). Existing techniques such as constrained RL, safe RL, and reward shaping, while effective, still depend on the transformation of the reward function, potentially complicating interpretability. Recently, policy regularization methods have been employed to achieve the desired behavior. One such method, known as affinity-based RL, has found applications in domains such as finance and machine ethics. In this paper, we introduce a variant called localized affinity-based RL (LAb-RL), which is versatile in state-specific decision-making. Our experiments show that agents can exhibit desired behaviors, and their actions in a given state can be interpreted through their localized affinities. We conclude by advocating the extension of this algorithm to other problems that necessitate state-specific and interpretable decision-making.