Post-hoc XAI Method for Visual Question Answering (VQA)
摘要
Visual Question Answering (VQA) systems, while advancing in intelligence, still face challenges in handling complex queries. Understanding the behavior of VQA models is crucial, especially in assessing their reliability in identifying relevant image regions for accurate responses. Although post-hoc explainable artificial intelligence (XAI) methods have demonstrated success in simpler tasks like image classification, their efficacy in more complex tasks like VQA remains unexplored. Moreover, the lack of a standardized evaluation method for the explanation heatmaps in VQA models raises questions about the trustworthiness of XAI methods. This paper proposes extending SIDU-XAI to VQA applications. It assesses visual explanations using the VQA-HAT dataset, which includes human-attention ground-truth maps, and the CLEVR-XAI dataset, comprising synthetic 3D scenes with objective ground-truth masks. Additionally, user studies were conducted to assess the impact of our post-hoc visual explanations on users’ understanding of model decisions. Our experiments led to the conclusion that the proposed SIDU-VQA method surpasses state-of-the-art RISE-VQA methods in both quantitative and qualitative aspects across three experiments. However, post-hoc XAI methods fall short of entirely fulfilling the complex standards of human expectations, which highlights the critical need for continued research into XAI techniques appropriate for VQA applications.