Robustness of Saliency-Based Explanations
摘要
This chapter covers core methods in robustness evaluation for saliency-based explanations. We begin by developing the fundamental principles and primary methods for computing saliency-based explanations and discuss the metrics that are used to evaluate their utility. We then argue that robustness is more than a metric; it is a core desideratum of saliency-based explanations without which they do more to erode trust than engender it. After formally defining different perspectives on robustness, we cover three methodologies for evaluating the robustness of saliency-based explanations: estimating robustness via sampling, auditing robustness with counter examples, and proving robustness with certification. The Chapter’s exposition is designed to give readers a solid technical understanding of the concepts at hand leaving wider literature discussion for the end of the chapter. We conclude with challenges and future directions for research in robustness of saliency-based explanations.