<p>Text-to-image person retrieval aims to identify the target person based on a given textual description query. The primary challenge is to learn the mapping of visual and textual modalities into a common latent space. Most existing methods rely on explicit local parts to simulate fine-grained correspondences between modalities, often lacking context information or introducing potential noise. Additionally, visual and textual features are separately extracted using pre-trained unimodal models, resulting in poor matching performance for multimodal data due to insufficient foundational alignment capabilities. To address these issues, we propose an innovative text-to-image person retrieval method, named text-to-image person retrieval with Implicit Relation Alignment and Contrastive Learning, IRACL for short, to combine implicit relationship alignment and contrastive learning. Specifically, the implicit relationship alignment module enhances the matching ability of multimodal data by mining hidden associations between multimodal data. Moreover, the contrastive learning module learns the similarity and dissimilarity between multimodal samples to extract more useful feature representations, thereby enhancing the semantic alignment capability. Our innovative framework demonstrates significant improvements over previous methods on three public datasets, marking notable progress in the task of visual-textual person retrieval.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Text-to-image person retrieval with implicit relation alignment and contrastive learning

  • Xiangyu Shui,
  • Zhenfang Zhu,
  • Yun Liu,
  • Hongli Pei,
  • Kefeng Li,
  • Huaxiang Zhang

摘要

Text-to-image person retrieval aims to identify the target person based on a given textual description query. The primary challenge is to learn the mapping of visual and textual modalities into a common latent space. Most existing methods rely on explicit local parts to simulate fine-grained correspondences between modalities, often lacking context information or introducing potential noise. Additionally, visual and textual features are separately extracted using pre-trained unimodal models, resulting in poor matching performance for multimodal data due to insufficient foundational alignment capabilities. To address these issues, we propose an innovative text-to-image person retrieval method, named text-to-image person retrieval with Implicit Relation Alignment and Contrastive Learning, IRACL for short, to combine implicit relationship alignment and contrastive learning. Specifically, the implicit relationship alignment module enhances the matching ability of multimodal data by mining hidden associations between multimodal data. Moreover, the contrastive learning module learns the similarity and dissimilarity between multimodal samples to extract more useful feature representations, thereby enhancing the semantic alignment capability. Our innovative framework demonstrates significant improvements over previous methods on three public datasets, marking notable progress in the task of visual-textual person retrieval.