A survey on explainable reinforcement learning: state of the art, challenges and opportunities
摘要
In recent years, the interest towards eXplainable AI (XAI) has been growing exponentially. Thanks to its capability of making the inner mechanisms of black-box models clearer, XAI practitioners and researchers have been drawing forward interpretability and explainability. This applies also for Reinforcement Learning (RL), where agents perform actions in a predefined environment and learn through their own experiences. It has indeed demonstrated to be versatile and successful in many domains due to its inherent learning adaptability. However, the intrinsic opacity of many RL decision-making processes, often obtained through deep neural networks, are not directly accessible or human-readable. This mechanism hinders the adoption of its methods and models as non-experts users usually must invest a high cognitive load to understand RL environments. To address this, the field of eXplainable RL (XRL) has been thriving recently and has developed its own study methodologies. This work provides an extensive and in-depth literature review of eXplainable Reinforcement Learning (XRL) to assess the current state of the art behind interpretability in RL and identify current opportunities and challenges. From this literature review, it becomes evident, in face of a series of preliminary and promising studies, the field of XRL still lacks a proper accounting towards full explainability. Moreover, it emerges that more effort should be devolved into developing paradigms with a human-in-the-loop factor and standardized metrics should be adopted to allow a fair comparison of different explanation methods.