Deep Foreground-Background Weighted Cross-modal Hashing
摘要
With the rapid growth of multi-modal data, deep cross-modal hashing algorithms provide a perfect solution for cross-modal retrieval tasks for their advantages of efficient retrieval speed and low storage consumption. Currently, the existing supervised cross-modal hashing methods, in order to efficiently extract structured information from raw data, generally gather on feature extraction of global information, however, all those methods ignore the weight differentiation between foreground and background information in a image. To address the issue, we propose a novel Deep Foreground-Background Weighted Cross-Modal Hashing(DFBWH) for supervised cross-modal retrieval. Specifically, the proposed method firstly performs target detection on the original image and select out candidate regions as target foreground entities. Then, the proposed method utilize the semantic interactions in the textual descriptions and tagging information as evaluation criteria, and use CLIP to detect the matching degree of the candidate regions. Eventually, under the supervision of the category labeling information, the hash loss function is utilized to obtain a high-quality hash code. Extensive experiments were carried out on two benchmark datasets, which demonstrate that DFBWH achieves better performance than the state-of-the-art baselines.