Cross-modal image-text retrieval is a challenging task due to the inherent ambiguity between modalities. However, most existing methods formulate this problem either with the coarse-grained information of the global image, ignoring the valuable fine-grained information implicit in local instances, or with the local features of the images and words, failing to provide a global understanding. In this paper, we propose a novel Fine-grained Feature Assisted Cross-modal Image-Text Retrieval (FiACR) model to learn a comprehensive and informative visual representation for cross-modal retrieval. Specifically, to address the absence of local information, we design a Local-Global Visual Features Fusion (LGVFF) module to aggregate global image and local instance information. By aggregation, FiACR can capture and leverage the images’ intricate details, which enables an accurate alignment between image and text. To enhance the global visual representation capability, we utilize the instance features to filter the global image feature’s attention and encourage it to focus on prominent regions in the image. Experimental results on several datasets show the competitive accuracy of our method compared to prior art.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Fine-grained Feature Assisted Cross-modal Image-text Retrieval

  • Chaofei Bu,
  • Xueliang Liu,
  • Zhen Huang,
  • Yuling Su,
  • Junfeng Tu,
  • Richang Hong

摘要

Cross-modal image-text retrieval is a challenging task due to the inherent ambiguity between modalities. However, most existing methods formulate this problem either with the coarse-grained information of the global image, ignoring the valuable fine-grained information implicit in local instances, or with the local features of the images and words, failing to provide a global understanding. In this paper, we propose a novel Fine-grained Feature Assisted Cross-modal Image-Text Retrieval (FiACR) model to learn a comprehensive and informative visual representation for cross-modal retrieval. Specifically, to address the absence of local information, we design a Local-Global Visual Features Fusion (LGVFF) module to aggregate global image and local instance information. By aggregation, FiACR can capture and leverage the images’ intricate details, which enables an accurate alignment between image and text. To enhance the global visual representation capability, we utilize the instance features to filter the global image feature’s attention and encourage it to focus on prominent regions in the image. Experimental results on several datasets show the competitive accuracy of our method compared to prior art.