The utilization of multi-spectral imaging, such as infrared, visible light, and ultraviolet, for recognizing defects in electrical equipment mostly focuses on static measurements and lacks exploration into the dynamic process of defect development. To better exploit dynamic measurements, this paper proposes a novel defect recognition method using tri-spectral videos. Specifically, a multi-modal spatio-temporal Transformer is presented to effectively decompose spatio-temporal features present in various modalities. Besides, a spatio-temporal multi-modal contrastive loss is introduced for self-supervised learning. By aligning extracted features both spatially and temporally across modalities, this loss helps mitigate confusion between modalities and improve the discriminative capacity of learned representations. To evaluate the proposed method, we self-collect a tri-spectral dataset, TROPED, which covers a wide range of dynamic defects in operational substation equipment, and benchmark results on the dataset. Experimental results demonstrate the effectiveness and robustness of the proposed method against other state-of-the-art methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-modal Spatio-temporal Transformer for Defect Recognition of Substation Equipment

  • Yiyang Yao,
  • Zexing Du,
  • Xue Wang,
  • Qing Wang

摘要

The utilization of multi-spectral imaging, such as infrared, visible light, and ultraviolet, for recognizing defects in electrical equipment mostly focuses on static measurements and lacks exploration into the dynamic process of defect development. To better exploit dynamic measurements, this paper proposes a novel defect recognition method using tri-spectral videos. Specifically, a multi-modal spatio-temporal Transformer is presented to effectively decompose spatio-temporal features present in various modalities. Besides, a spatio-temporal multi-modal contrastive loss is introduced for self-supervised learning. By aligning extracted features both spatially and temporally across modalities, this loss helps mitigate confusion between modalities and improve the discriminative capacity of learned representations. To evaluate the proposed method, we self-collect a tri-spectral dataset, TROPED, which covers a wide range of dynamic defects in operational substation equipment, and benchmark results on the dataset. Experimental results demonstrate the effectiveness and robustness of the proposed method against other state-of-the-art methods.