CCUH: CLIP-Based Clustering Method for Unsupervised Hashing Multi-modal Retrieval
摘要
Unsupervised methods have recently garnered significant attention owing to their effective retrieval capabilities and minimal storage requirements. However, current unsupervised techniques do not sufficiently capture the information about the joint occurrence of data and primarily rely on high-dimensional features to construct instance similarity matrices. This approach fails to effectively guide the learning of hash codes. To tackle the previously mentioned challenges, we put forward CLIP-Based Clustering Method for Unsupervised Hashing Multi-Modal Retrieval(CCUH). First, we extract pre-existing image and text features from the CLIP model to serve as the input data for the neural network. Then, we design a novel multimodal association matrix generator that employs cosine similarity and KMeans clustering algorithms. This generator combines the high-dimensional features obtained by the network with discriminative embeddings to build a cross-channel semantic association matrix enriched with additional information. This approach effectively supervises the learning of hash codes and improves the arrangement of semantic information across various modalities. Extensive empirical findings Indicate that on two commonly utilized datasets, our proposed CCUH method surpasses existing leading methods in cross-modal search tasks.