The modality gap between visible and infrared images poses significant challenges in visible-infrared person re-identification, as essential color and lighting information is often lost in infrared images. Existing methods mainly focus on learning global representations of the entire person, overlooking fine-grained variations in shape and texture across different body parts. To address this, we propose a Part Feature Mining Network (PFMN) that extracts discriminative features from individual body parts and aligns them at a fine-grained level to mitigate cross-modality discrepancies. To capture fine-grained details, we first propose a Frequency-domain Detail Guidance (FDG) module, which leverages Fast Fourier Transform to highlight frequency-domain components that reveal subtle texture and shape information. Then, we develop a Feature Expansion Module (FEM) that enriches features by combining global and local contexts via dilated and focal convolutions. Furthermore, we design a Part Feature Relation Mining (PFRM) module to extract horizontal part-level features and explore correlations among body parts. Additionally, we introduce a Cross-Domain Alignment (CDA) loss to improve part-level feature alignment and a Hardness Focus Triplet (HFT) loss to emphasize informative hard samples. Extensive experiments on SYSU-MM01, RegDB, and LLCM demonstrate that PFMN outperforms state-of-the-art methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Frequency-Enhanced Part Feature Mining and Cross-Modality Alignment for Visible-Infrared Person Re-Identification

  • Yuqing Wu,
  • Yongkang Ding,
  • Liyan Zhang

摘要

The modality gap between visible and infrared images poses significant challenges in visible-infrared person re-identification, as essential color and lighting information is often lost in infrared images. Existing methods mainly focus on learning global representations of the entire person, overlooking fine-grained variations in shape and texture across different body parts. To address this, we propose a Part Feature Mining Network (PFMN) that extracts discriminative features from individual body parts and aligns them at a fine-grained level to mitigate cross-modality discrepancies. To capture fine-grained details, we first propose a Frequency-domain Detail Guidance (FDG) module, which leverages Fast Fourier Transform to highlight frequency-domain components that reveal subtle texture and shape information. Then, we develop a Feature Expansion Module (FEM) that enriches features by combining global and local contexts via dilated and focal convolutions. Furthermore, we design a Part Feature Relation Mining (PFRM) module to extract horizontal part-level features and explore correlations among body parts. Additionally, we introduce a Cross-Domain Alignment (CDA) loss to improve part-level feature alignment and a Hardness Focus Triplet (HFT) loss to emphasize informative hard samples. Extensive experiments on SYSU-MM01, RegDB, and LLCM demonstrate that PFMN outperforms state-of-the-art methods.