Traditional cross-modal retrieval (CMR) methods are usually based on the assumption that all categories of retrieval samples are included in the training set, making it hard to apply for CMR of samples with new categories. As a result, zero-shot cross-modal retrieval (ZS-CMR) emerges with increasing research interest. However, existing ZS-CMR methods cannot fully transfer the semantic knowledge from seen to unseen classes due to the weak semantic association between visible and invisible categories. In this study, this paper proposes a semantic cross-self-reconstruction with graph convolutional network (SCSRGCN) for ZS-CMR. Specifically, it employs auto-encoders to convert heterogeneous multiple modalities into common space while retaining the modality-specific information. Then, it applies GCN to exploit the intra-modal relationship, which can guide the relationship learning of another modal encoder. Finally, it develops a semantic similarity transfer module to enhance further the zero-shot feature learning of heterogeneous modalities for ZS-CMR. Experimental results on three widely used datasets show that the proposed SCSRGCN outperforms the state-of-the-art methods on both ZS-CMR and CMR.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Semantic Cross-Self-Reconstruction with Graph Convolutional Network for Zero-Shot Cross-Modal Retrieval

  • Longfa Liu,
  • Kexin Gao,
  • Chuang Li,
  • Imad Rida,
  • Shaohua Teng,
  • Lunke Fei

摘要

Traditional cross-modal retrieval (CMR) methods are usually based on the assumption that all categories of retrieval samples are included in the training set, making it hard to apply for CMR of samples with new categories. As a result, zero-shot cross-modal retrieval (ZS-CMR) emerges with increasing research interest. However, existing ZS-CMR methods cannot fully transfer the semantic knowledge from seen to unseen classes due to the weak semantic association between visible and invisible categories. In this study, this paper proposes a semantic cross-self-reconstruction with graph convolutional network (SCSRGCN) for ZS-CMR. Specifically, it employs auto-encoders to convert heterogeneous multiple modalities into common space while retaining the modality-specific information. Then, it applies GCN to exploit the intra-modal relationship, which can guide the relationship learning of another modal encoder. Finally, it develops a semantic similarity transfer module to enhance further the zero-shot feature learning of heterogeneous modalities for ZS-CMR. Experimental results on three widely used datasets show that the proposed SCSRGCN outperforms the state-of-the-art methods on both ZS-CMR and CMR.