A CNN-transformer hybrid model and a multi-modal multi-stage training strategy for visible-infrared person re-identification
摘要
Visible-Infrared Person Re-Identification (VI Re-ID) is designed to match visible pedestrian images with infrared pedestrian images. The modality discrepancy between visible and infrared images is the biggest challenge for VI Re-ID. Existing VI Re-ID approaches mostly weaken the effect of modality discrepancy by extracting discriminative features, while ignoring the ability of models to adapt to modality variations. To address this issue, we construct a CNN-Transformer hybrid model (CTHM) for VI Re-ID, which mainly consists of a dual-stream ResNet-50 network, a dual-stream VIT network, and a multi-modal feature fusion module (MFFM). In particular, we design a multi-modal multi-stage training strategy (MMTS). Given visible images and infrared images, MMTS first uses Cycle GAN to generate fake infrared images and fake visible images. Then, MMTS sequentially employs visible images and fake visible images, infrared images and fake infrared images, fake visible images and fake infrared images, and visible images and infrared images, to train CTHM in stages in order to progressively improve its adaptive ability to modality variations, thus ensuring that CTHM can obtain more discriminative modality-invariant features and better attenuates the influence of modality discrepancy. We conduct extensive experiments on two benchmark datasets, SYSU-MM01 and RegDB, and the results show that our method reaches the current advanced level.