<p>Text-image person re-identification (TIReID) is a cross-modal retrieval task that aims to query person images with corresponding identities through natural language descriptions. The key to this task is the effective alignment of cross-modal features between image-text pairs. Many methods have achieved promising experimental results by fine-tuning pre-trained visual language models. However, existing methods largely depend on accurate and high-quality text annotations, which require substantial time and resources. During dataset construction, this reliance results in incorrect sample pair matching and introduces coupled noisy labels in TIReID. Although some prior studies have achieved relatively robust outcomes in addressing the noise correspondence problem in TIReID, they still face several challenges: (1) Lacking spatial detail of person images: Previous research predominantly utilizes pre-trained vision transformers for visual feature extraction. However, simply relying on position encoding in ViT can result in insufficient learning of the spatial structural characteristics of pedestrians, thereby limiting the effectiveness of the model in distinguishing subtle identity-related variations. (2) Neglect of identity category relationships: Many prior approaches primarily address noise relationships using sample-level similarity and loss responses, often overlooking the predicted relationships between identity categories. These relationships are crucial for guiding the model to focus on shared identity characteristics. To address these challenges, we propose Spatial Enhanced Multi-Level Alignment Learning (SE-MLAL). SE-MLAL includes a consistent noise detection module that predicts the correctness of sample pairs’ correspondence at both the sample and class level, leveraging the consistency between these predictions to achieve accurate dataset division. Building upon this foundation, we employ sample-wise triplet loss and class-wise alignment loss to facilitate hierarchical feature alignment loss. Experimental results across three datasets substantiate the effectiveness and robustness of our method.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Spatial enhanced multi-level alignment learning for text-image person re-identification with coupled noisy labels

  • Jiacheng Zhao,
  • Haojie Che,
  • Yongxi Li

摘要

Text-image person re-identification (TIReID) is a cross-modal retrieval task that aims to query person images with corresponding identities through natural language descriptions. The key to this task is the effective alignment of cross-modal features between image-text pairs. Many methods have achieved promising experimental results by fine-tuning pre-trained visual language models. However, existing methods largely depend on accurate and high-quality text annotations, which require substantial time and resources. During dataset construction, this reliance results in incorrect sample pair matching and introduces coupled noisy labels in TIReID. Although some prior studies have achieved relatively robust outcomes in addressing the noise correspondence problem in TIReID, they still face several challenges: (1) Lacking spatial detail of person images: Previous research predominantly utilizes pre-trained vision transformers for visual feature extraction. However, simply relying on position encoding in ViT can result in insufficient learning of the spatial structural characteristics of pedestrians, thereby limiting the effectiveness of the model in distinguishing subtle identity-related variations. (2) Neglect of identity category relationships: Many prior approaches primarily address noise relationships using sample-level similarity and loss responses, often overlooking the predicted relationships between identity categories. These relationships are crucial for guiding the model to focus on shared identity characteristics. To address these challenges, we propose Spatial Enhanced Multi-Level Alignment Learning (SE-MLAL). SE-MLAL includes a consistent noise detection module that predicts the correctness of sample pairs’ correspondence at both the sample and class level, leveraging the consistency between these predictions to achieve accurate dataset division. Building upon this foundation, we employ sample-wise triplet loss and class-wise alignment loss to facilitate hierarchical feature alignment loss. Experimental results across three datasets substantiate the effectiveness and robustness of our method.