Recently, the development and integration of Artificial Intelligence (AI) has accelerated and been popularized widely throughout modern society. AI is becoming a powerful tool ranging from leisurely use to critical applications. However, due to the black-box nature of some AI approaches such as Deep Reinforcement Learning (DRL), complex AI algorithms now face growing concerns of trust in ethical and responsible decision-making. EXplainable Artificial Intelligence (XAI) is a subfield of AI focused on deriving interpretable information from incomprehensible statistics to generate explanations for an AI’s decisions. This paper proposes an architecture that combines 2 XAI techniques, Testable Concept Activation Vectors (TCAV) and Reward Decomposition, to create goal-oriented explanations. The XAI approach is tested in a simulated movement prediction environment where a DRL agent is trained to represent different human concepts and goal prioritizations; we can confidently distinguish those concepts between agents in a human-centric framework. Results obtained demonstrate our method allows users to insert their own high-level thinking into XAI and use it to generate explanations.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Human Decision-Making Concepts with Goal-Oriented Reasoning for Explainable Deep Reinforcement Learning

  • Chris Lee,
  • Eduardo Benitez Sandoval,
  • Francisco Cruz

摘要

Recently, the development and integration of Artificial Intelligence (AI) has accelerated and been popularized widely throughout modern society. AI is becoming a powerful tool ranging from leisurely use to critical applications. However, due to the black-box nature of some AI approaches such as Deep Reinforcement Learning (DRL), complex AI algorithms now face growing concerns of trust in ethical and responsible decision-making. EXplainable Artificial Intelligence (XAI) is a subfield of AI focused on deriving interpretable information from incomprehensible statistics to generate explanations for an AI’s decisions. This paper proposes an architecture that combines 2 XAI techniques, Testable Concept Activation Vectors (TCAV) and Reward Decomposition, to create goal-oriented explanations. The XAI approach is tested in a simulated movement prediction environment where a DRL agent is trained to represent different human concepts and goal prioritizations; we can confidently distinguish those concepts between agents in a human-centric framework. Results obtained demonstrate our method allows users to insert their own high-level thinking into XAI and use it to generate explanations.