Text-based Person Search (TPS) aims to retrieve the most matched target in person image gallery through given text description, yet faces the challenge of semantic misalignment due to text with ambiguously modifying and diverse images. In this paper, we introduce the Dependency-aware and Semantic Alignment Network (DSAN) for extracting correspondences entity and attribute. Specifically, we design a dependency-driven syntactic procedure to generate a dependency mask for each text descriptions. Dependency mask uses dependency distance as an initial heuristic cue to measure the importance between entities and attributes. Our Dependency Intervention Self-Attention (DISA) module incorporates dependency mask to intervene the attention of entities and attributes in syntactic contexts. On local scale, an additional Part-Channel Attention (PCA) module is introduced to adaptively extract locally corresponding text and visual features, and a shared Near Correlation Fusion (NCF) module is used to enhance cross-modal interaction. We introduced DISA into existing advanced baselines, and extensive experiments on CUHK-PEDES and ICFG-PEDES datasets demonstrated the effectiveness of our method, with up to 0.61% and 0.66% improvement on Rank-1, respectively.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Entity Attribute Correspondence Learning for Text-Based Person Search

  • Wei Xia,
  • Jianguo Chen,
  • Wenguang Gan,
  • Shaomin Xie,
  • Xinpan Yuan

摘要

Text-based Person Search (TPS) aims to retrieve the most matched target in person image gallery through given text description, yet faces the challenge of semantic misalignment due to text with ambiguously modifying and diverse images. In this paper, we introduce the Dependency-aware and Semantic Alignment Network (DSAN) for extracting correspondences entity and attribute. Specifically, we design a dependency-driven syntactic procedure to generate a dependency mask for each text descriptions. Dependency mask uses dependency distance as an initial heuristic cue to measure the importance between entities and attributes. Our Dependency Intervention Self-Attention (DISA) module incorporates dependency mask to intervene the attention of entities and attributes in syntactic contexts. On local scale, an additional Part-Channel Attention (PCA) module is introduced to adaptively extract locally corresponding text and visual features, and a shared Near Correlation Fusion (NCF) module is used to enhance cross-modal interaction. We introduced DISA into existing advanced baselines, and extensive experiments on CUHK-PEDES and ICFG-PEDES datasets demonstrated the effectiveness of our method, with up to 0.61% and 0.66% improvement on Rank-1, respectively.