WFSFA-Net: Weighted Feature Supplementation and Cross-Modal Feature Alignment for Visible-Infrared Person Re-identification
摘要
Visible-infrared person re-identification (VI-ReID) has captured growing attention for its applications in surveillance within low-light environments. Due to the substantial modality discrepancy and pedestrian variations, VI-ReID remains a challenging task. In this paper, a weighted feature supplementation and feature alignment network (WFSFA-Net) is presented to tackle the primary challenges in VI-ReID. The proposed approach consists of two modules - the Weighted Feature Supplementation (WFS) module and the Cross-modal Feature Alignment (CMFA) module. WFS module can generate supplementary embeddings to mine informative representations to narrow the modality gap. CMFA module mines structural relationships between multi-modal features of the same pedestrian and then aligns these features of the two modalities by using the shortest path algorithm. This process can enhance the model’s robustness and generalization against pedestrian variations. Extensive experiments conducted on the SYSU-MM01 and RegDB datasets demonstrate the effectiveness of our approach, outperforming state-of-the-art methods by more than 3% in accuracy.