Learning Multi-semantic Based on Cross-Attention for Image-Text Retrieval
摘要
Cross-modal retrieval is a challenging problem for processing multi-modal information, primarily addressing the semantic gap between different modalities. However, current research mainly focuses on aligning image and text regions, which often ignore the importance of contextual background information. Specially, there are uncertainties among heterogeneous modalities which persist in understanding semantic information of cross-modal retrieval. In this paper, we propose a novel Multi-Semantic Based on Cross-Attention Network (MSCA) for cross-modal retrieval tasks. Specifically, a cross-extraction of contextual semantic feature encoder module and a graph structure enhancement module are firstly designed to more effectively address the issue of neglecting contextual background information and to further enhance the fusion of information and features, respectively. Secondly, we effectively solve the cross-modal alignment problem by employing a probability distribution encoder, simultaneously, we integrate three pre-training models, which include image-text contrastive learning, image-text matching loss and masked language model loss for reducing uncertainties among different modalities. We conducted extensive experiments on two large datasets, MS-COCO and Flickr30K. The experimental results show that the MSCA model outperforms other existing methods on both datasets, demonstrating its effectiveness and significant performance advantages.