Interpretable capsule networks via self attention routing on spatially invariant feature surfaces
摘要
The accurate and efficient evaluation and classification of situational images is fundamental to making informed and effective decisions. However, current classification approaches based on convolutional neural networks often suffer from limited generalization and robustness, particularly when processing data characterized by abstract class features and pronounced spatial attributes. Additionally, the “black-box” nature of deep neural network architectures poses significant challenges to their application in fields with stringent security requirements. To address these limitations, this paper introduces a novel Spatially Invariant Self-Attention Capsule Network (SISA-CapsNet), designed to encode interpretable spatial features for classification tasks. SISA-CapsNet employs capsules to encode spatial features from specific image regions and classifies these features through a self-attention routing mechanism. Specifically, spatially invariant feature surfaces with dimensions identical to the input image are generated and stacked to form feature capsules, each encoding spatial features from distinct regions. The self-attention mechanism calculates coupling coefficients, clustering feature capsules into class capsules. This architecture integrates a spatially invariant feature extraction structure, facilitating pixel-level encoding of regional spatial features, and leverages self-attention to effectively capture the relative importance of different spatial regions for classification. Together, these two mechanisms constitute an interpretable classification framework. Experimental validation on benchmark datasets and battlefield situational image datasets with pronounced spatial characteristics demonstrates that the proposed method not only achieves superior classification performance but also offers interpretability closely aligned with human cognitive processes. Furthermore, comparative analyses with existing visual interpretability method underscore the enhanced interpretability of SISA-CapsNet.