Modality Confusion Learning: A Versatile Framework for Visible-Infrared Re-identification
摘要
Due to the discrepancy between visible and infrared modalities, most existing biometric recognition tasks like Person Re-Identification (ReID) suffered from obtaining modality-irrelevant and identity-relevant features between different modalities. In this work, we propose an end-to-end Modality Confusion Learning Network+ (MCLNet+), which ensures the extracted features are related to the identity and reliable beyond the different modalities. Unlike previous methods of designing complicated branches to learn modality-invariant features, MCLNet+ learns modality invariance and identity relevance features in a concise single framework. We design a Modality Confusion Learning core (MCL core) to further conduct the discriminative analysis between visible and infrared features and exploit a max-min game to decouple the properties of different modalities in the representation space. In addition, we propose the prior-aware marginal center aggregation strategy to explore the latent correspondence between cameras and identities, which leads the model to learn the camera-awareness and identity-awareness. Our MCLNet+ outperforms the existing state-of-the-art performances on both visible-infrared recognition tasks. On the large-scale SYSU-MM01 dataset, our model can achieve 71.10 % and 65.78 % in terms of Rank-1 accuracy and mAP scores.