Human Decision-Making Concepts with Goal-Oriented Reasoning for Explainable Deep Reinforcement Learning
摘要
Recently, the development and integration of Artificial Intelligence (AI) has accelerated and been popularized widely throughout modern society. AI is becoming a powerful tool ranging from leisurely use to critical applications. However, due to the black-box nature of some AI approaches such as Deep Reinforcement Learning (DRL), complex AI algorithms now face growing concerns of trust in ethical and responsible decision-making. EXplainable Artificial Intelligence (XAI) is a subfield of AI focused on deriving interpretable information from incomprehensible statistics to generate explanations for an AI’s decisions. This paper proposes an architecture that combines 2 XAI techniques, Testable Concept Activation Vectors (TCAV) and Reward Decomposition, to create goal-oriented explanations. The XAI approach is tested in a simulated movement prediction environment where a DRL agent is trained to represent different human concepts and goal prioritizations; we can confidently distinguish those concepts between agents in a human-centric framework. Results obtained demonstrate our method allows users to insert their own high-level thinking into XAI and use it to generate explanations.