<p>While detecting a decline in collective performance is often straightforward, localizing the root cause can be significantly more challenging—particularly in systems characterized by dynamic relationships and complex interdependencies. To address this challenge, several studies have attempted to utilize explainable AI (XAI) techniques to localize causative agents. However, these approaches have not achieved sufficient accuracy in root-cause localization. This study proposes a method that integrates the Vision Transformer (ViT) with Gradient Self-Attention Maps (Grad-SAM)—an XAI technique—to localize anomalous agents within a swarm system. To evaluate the effectiveness of the proposed approach, we conduct comparative experiments against conventional convolutional neural networks (CNNs) combined with Gradient-weighted Class Activation Mapping (Grad-CAM), using the Boid model as the swarm simulation environment. Our comparative experiments against CNN/Grad-CAM using the Boid model revealed that anomaly detection and cause localization are a dual challenge with different requirements for model architecture. While CNNs demonstrated superiority in anomaly detection, particularly in scenarios with rare anomalies, ViT showed a clear advantage in cause localization. Specifically, ViT’s Grad-SAM generates higher-resolution and more stable saliency maps than Grad-CAM, thereby achieving high robustness with minimal performance degradation under a practical fixed-threshold setting. Furthermore, it outperformed CNNs under more complex situations involving mixed anomalies and background noise, quantitatively demonstrating that the proposed method can be a more practical and reliable technique for cause localization.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ViT-XAI: vision transformer and XAI enable robust detection and cause localization in swarm anomalies

  • Yohei Fukuyama,
  • Masao Kubo,
  • Hiroshi Sato

摘要

While detecting a decline in collective performance is often straightforward, localizing the root cause can be significantly more challenging—particularly in systems characterized by dynamic relationships and complex interdependencies. To address this challenge, several studies have attempted to utilize explainable AI (XAI) techniques to localize causative agents. However, these approaches have not achieved sufficient accuracy in root-cause localization. This study proposes a method that integrates the Vision Transformer (ViT) with Gradient Self-Attention Maps (Grad-SAM)—an XAI technique—to localize anomalous agents within a swarm system. To evaluate the effectiveness of the proposed approach, we conduct comparative experiments against conventional convolutional neural networks (CNNs) combined with Gradient-weighted Class Activation Mapping (Grad-CAM), using the Boid model as the swarm simulation environment. Our comparative experiments against CNN/Grad-CAM using the Boid model revealed that anomaly detection and cause localization are a dual challenge with different requirements for model architecture. While CNNs demonstrated superiority in anomaly detection, particularly in scenarios with rare anomalies, ViT showed a clear advantage in cause localization. Specifically, ViT’s Grad-SAM generates higher-resolution and more stable saliency maps than Grad-CAM, thereby achieving high robustness with minimal performance degradation under a practical fixed-threshold setting. Furthermore, it outperformed CNNs under more complex situations involving mixed anomalies and background noise, quantitatively demonstrating that the proposed method can be a more practical and reliable technique for cause localization.