Zero-Shot Referring Image Segmentation with Hierarchical Prompts and Frequency Domain Fusion
摘要
Zero-shot referring image segmentation is a method of accurately identifying masks that are most relevant to referring expressions, without relying on pixel-level annotations. This process involves the generation of masks and the matching of these masks to text, which are essential for producing precise, high-quality masks and exploring visual-textual relationships. Traditional methods often fail to generate sufficiently detailed masks and contain excessive redundancies. To address these issues, we introduce a hierarchical prompts mask generation network that significantly improves masks quality and reduces redundancy. Furthermore, we address the limitations of spatial sensitivity and detail recognition in the CLIP model through a method that exploits spatial orientation descriptions in textual cues to extract visual feature. This guides CLIP visual encoder to focus on objects in a specific space within the image. Moreover, the mere addition of visual and text feature does not fully exploit contour information, edge feature in images and crucial information in text. Consequently, we utilize the Non-subsampled Contourlet Transform and the Haar Wavelet Transform to decompose and fuse visual and text feature separately in the frequency domain to fully utilize feature information. Our method has been extensively tested on the RefCOCO, RefCOCO+, and RefCOCOg datasets, demonstrating significant improvements over existing zero-shot referring image segmentation methods.