Ranking-Aware Uncertainty for Text-Guided Image Retrieval
摘要
Text-guided image retrieval is to incorporate conditional text to better capture users’ intent. Traditionally, existing methods mainly focus on minimizing the embedding distances between the source inputs and the targeted images, using the provided triplets source image, source text, target image. However, such triplet optimization may limit the learned retrieval model to capture uncertain ranking information caused by the semantic diversity among texts and images. For example, “I want nicer clothes” where “nicer” may cause semantic uncertainty. To alleviate this problem, we propose a novel ranking-aware uncertainty approach for text-guided image retrieval. Firstly, multiple uncertainty augmenters are proposed to capture uncertain ranking information of the multimodal data. We use Gaussian distribution derived from both combined and target features to express richer ranking relationships. Then, we further propose a ranking-aware uncertainty loss to mine the ranking uncertainty among samples. The top probabilities for other uncertain images are also maximized. Finally, a distribution regularization is introduced to better align the distributional representations of source inputs and targeted images. Compared to the existing state-of-the-art methods, our proposed method achieves significant results on two public datasets for text-guided image retrieval.