Visible-Infrared Person Re-identification (VI-ReID) is a challenging task that involves matching visible and infrared person images across multiple camera views. The huge gap between visible and infrared modalities has become a significant bottleneck. Existing works typically employ dual-stream networks to extract shared modality representation, yet struggle to relieve such gap, resulting in inferior performance. To overcome this issue, we propose a robust synthetic modality learning with a dual-level alignment method (RSDL) for VI-ReID that aims to generate a robust synthetic modality as a bridge to guide cross-modality alignment. Specifically, the hetero-modality fusion (HMF) strategy is introduced to generate the robust synthetic modality by using multi-scale feature fusion with a structure rebuild module (SRM) and a cross-modality spatial alignment (CSA) module. The strategy incorporates rich semantic structural patterns from visible and infrared images to handle the modality variation. Additionally, we design the dual-level regulation loss to jointly explore the stable feature relationships among three modalities at both the instance and distribution levels for cross-modality alignment. This facilitates discovering modality-consistent and identity-aware representations. Extensive experiments on three VI-ReID benchmarks demonstrate the effectiveness of our proposed method.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Learning a Robust Synthetic Modality with Dual-Level Alignment for Visible-Infrared Person Re-identification

  • Zichun Wang,
  • Xu Cheng

摘要

Visible-Infrared Person Re-identification (VI-ReID) is a challenging task that involves matching visible and infrared person images across multiple camera views. The huge gap between visible and infrared modalities has become a significant bottleneck. Existing works typically employ dual-stream networks to extract shared modality representation, yet struggle to relieve such gap, resulting in inferior performance. To overcome this issue, we propose a robust synthetic modality learning with a dual-level alignment method (RSDL) for VI-ReID that aims to generate a robust synthetic modality as a bridge to guide cross-modality alignment. Specifically, the hetero-modality fusion (HMF) strategy is introduced to generate the robust synthetic modality by using multi-scale feature fusion with a structure rebuild module (SRM) and a cross-modality spatial alignment (CSA) module. The strategy incorporates rich semantic structural patterns from visible and infrared images to handle the modality variation. Additionally, we design the dual-level regulation loss to jointly explore the stable feature relationships among three modalities at both the instance and distribution levels for cross-modality alignment. This facilitates discovering modality-consistent and identity-aware representations. Extensive experiments on three VI-ReID benchmarks demonstrate the effectiveness of our proposed method.