Existing cross-modal hashing methods have made progress in enhancing retrieval capabilities and reducing model size, but they struggle to balance retrieval performance across different channels, leading to increased robustness.These methods often show low integration of multi-channel semantic information and fail to address image-text heterogeneity balance, focusing solely on enhancing retrieval accuracy, which can lead to high model robustness issues. We propose the Joint Modal Heterogeneous Balance Hashing for Unsupervised Cross-Modal Retrieval (JMBH) to address this. We utilise the large model CLIP to process raw data, facilitating multi-channel semantic integration. We then design multi-channel fusion modalities to explore co-occurrence information across channels and develop intra- and inter-channel constraints to mine this information. Extensive experiments on three datasets validate JMBH’s effectiveness in balancing image-text heterogeneity and reducing robustness.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Joint Modal Heterogeneous Balance Hashing for Unsupervised Cross-Modal Retrieval

  • Jie Zhang,
  • Mingyong Li

摘要

Existing cross-modal hashing methods have made progress in enhancing retrieval capabilities and reducing model size, but they struggle to balance retrieval performance across different channels, leading to increased robustness.These methods often show low integration of multi-channel semantic information and fail to address image-text heterogeneity balance, focusing solely on enhancing retrieval accuracy, which can lead to high model robustness issues. We propose the Joint Modal Heterogeneous Balance Hashing for Unsupervised Cross-Modal Retrieval (JMBH) to address this. We utilise the large model CLIP to process raw data, facilitating multi-channel semantic integration. We then design multi-channel fusion modalities to explore co-occurrence information across channels and develop intra- and inter-channel constraints to mine this information. Extensive experiments on three datasets validate JMBH’s effectiveness in balancing image-text heterogeneity and reducing robustness.