Weighted semantic feature based self-supervised deep cross-modal hashing
摘要
To fast respond the large-scale cross-modal retrieval task, the deep cross-modal hashing algorithms map different modalities into the low-dimensional Hamming space and measure their similarity degree using Hamming distance. The self-supervised methods train deep hashing network based on the semantic multi-label and preserve the original semantic relationship in the Hamming space. However, most self-supervised methods neglect the semantic weight inconsistency problem among the image, text and multi-label. Moreover, they ignore preserving the relative ranking orders among the retrieval samples. To address the above issues, we propose a novel method termed weighted semantic feature based self-supervised deep cross-modal hashing (WFSCH). Firstly, we design the weighted semantic feature module to adaptively generate the weighted semantic features for the image, text and multi-label. The weighted semantic features improve the accuracy of describing different modal content and enhance the self-supervised semantic constraint. Secondly, to further improve the cross-modal retrieval performance, we preserve the original intra- and inter-modal triplet ranking relationship in the Hamming space. Finally, to minimize the discrepancy between the pairwise hashing similarity and the multi-label semantic relationship, we design the multi-label semantic similarity preserving loss based on the positive-constraint Kullback-Leibler (KL) divergence, which fully explores different modal multi-label semantic information. We conduct extensive experiments on four publicly available datasets including MIRFLICKR-25K, NUS-WIDE, MS COCO2014 and IAPR TC-12. The experimental results demonstrate that the proposed WFSCH algorithm outperforms the state-of-the-art cross-modal hashing algorithms.