Image-text retrieval (ITR) has made significant progress in recent years. However, it still faces two major challenges. The first challenge is the problem of intra-modal semantic loss issue, which is related to the lack of semantic associations between single-modal data. The second challenge is that different modal data cannot be effectively mapped to the same shared space, resulting in inconsistent representations between multimodal data, making it difficult to perform effective alignment and fusion. These challenges lead to limitations in the generality and retrieval accuracy of existing ITR retrieval models and challenges in the validity and reliability of these methods in practical applications. We propose two new methods to address these challenges: the Unimodal Momentum Soft Label Alignment (UMSA) method and the Multimodal Data Potential Projection (MMLP) method. Our methods aim to establish semantic links between unimodal data and overcome the underfitting problem of linearly mapping multimodal data into the same shared space. Our method has been extensively experimentally validated on various ITR models and datasets, all showing significant improvements in retrieval performance, including zero-sample retrieval performance. In addition, this approach is compatible with a wide range of ITR retrieval models, thereby improving model generality and accuracy.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Momentum-Based Uni-modal Soft-Label Alignment and Multi-modal Latent Projection Networks for Optimizing Image-Text Retrieval

  • Xiaole Zhu,
  • Zongtao Duan,
  • Junchen Huang,
  • Xing Sheng

摘要

Image-text retrieval (ITR) has made significant progress in recent years. However, it still faces two major challenges. The first challenge is the problem of intra-modal semantic loss issue, which is related to the lack of semantic associations between single-modal data. The second challenge is that different modal data cannot be effectively mapped to the same shared space, resulting in inconsistent representations between multimodal data, making it difficult to perform effective alignment and fusion. These challenges lead to limitations in the generality and retrieval accuracy of existing ITR retrieval models and challenges in the validity and reliability of these methods in practical applications. We propose two new methods to address these challenges: the Unimodal Momentum Soft Label Alignment (UMSA) method and the Multimodal Data Potential Projection (MMLP) method. Our methods aim to establish semantic links between unimodal data and overcome the underfitting problem of linearly mapping multimodal data into the same shared space. Our method has been extensively experimentally validated on various ITR models and datasets, all showing significant improvements in retrieval performance, including zero-sample retrieval performance. In addition, this approach is compatible with a wide range of ITR retrieval models, thereby improving model generality and accuracy.