Cross Modality Fusion Network with Feature Alignment and Salient Object Exchange for Single Image 3D Shape Retrieval
摘要
The image-based 3D shape retrieval (IBSR) aims to retrieve 3D shapes that are similar to the query image. Most methods consider metric learning, which involves mapping images and 3D shapes to a low-dimensional space. This enables greater similarity between images and 3D shapes of the same instance, while images and 3D shapes of different instances are dissimilar. However, most existing methods do not consider the fusion of information across modalities. By leveraging complementary knowledge contained in different modalities, integrating data from different modalities into a single representation comprehensively represents the data, which enhances the data representation capability and thus facilitates retrieval. Therefore we propose a new method that takes into account information across different modalities. Firstly, we introduce a cross modality fusion network. The cross modality fusion network is primarily an attention mechanism network. By employing this attention mechanism network to fuse modal information, the network can determine the probability of similarity between the input query image and 3D shape. Secondly, to alleviate the difficulty of modal fusion, we propose a feature alignment module based on contrastive learning. This module includes instance discrimination and cross domain feature alignment modules, which align features before modal fusion. Finally, we propose salient object exchange, which further assists in modal fusion. Experiments on three commonly used datasets, i.e., Pix3D, Stanford Cars, and Comp Cars, demonstrates the effectiveness of the proposed method.