A Minimalistic Definition of XAI Explanations
摘要
In the field of XAI, the concept of explanation clearly has a central role. But what fundamentally constitutes an XAI explanation is far from settled. Many arguments exist around what qualitative properties XAI explanations should have, leaning heavily into their relevance to the real world and normative preferences derived thereof. Here, we argue instead that such explanations ultimately originate from a model world and limit our definition to those. Therefore, real-world considerations such as fairness, ethics, fidelity and trustworthiness are absent, leaving, as we believe, a more pristine environment in which to investigate ideal XAI explanations. Stripped down to the essentials of the model world, there is only the AI model relating its inputs and outputs. Assuming further, that the process of identifying an explanation involves backtracking from the model outputs to its inputs, we are left with a very clear view of what ideally constitutes an explanation in a model world: an explanation is a solution of an inverse problem posed to an AI model. Thus even before any real world considerations can take place, we note that for complex AI models inverse problems are sufficiently hard to be essentially unsolvable. Therefore, such ideal explanations are unattainable and XAI methods must make compromises in accuracy, completeness and reliability already in order to be practical.