Visual-Language Pretraining-Driven Zero-Shot Sketch-Based Image Retrieval
摘要
The ZS-SBIR task that aims to address the inherent heterogeneity problem between sketches and images, as well as the knowledge transfer problem. However, most existing works utilize models pretraining on a single modality, leading to an inability to effectively establish relationships between images and text, and preventing sufficient knowledge transfer. In this paper, we rethink and present a visual-language pretraining-driven model (termed VLPD) for ZS-SBIR task. Specifically, we leverage the CLIP model to extract features for sketches, images, and text, to achieve knowledge transfer. Meanwhile, we propose a semantic consistency enhancement method that exploits textual modalities to enhance semantic consistency between sketches and images, alleviating the inherent heterogeneity trouble. Extensive experiments on three ZS-SBIR datasets reveal that VLPD dramatically outperforms the SOTA single-modal pretrained models.