The performance of existing cross-modal matching techniques highly depends on the precise alignment of multimodal data. Nevertheless, in the process of data collection or annotation, certain data pairs may suffer from compromised correspondences, giving rise to the problem of noisy correspondence. To tackle this hurdle, we propose an innovative approach, namely Neighbor-Replacing Cross-modal (NeiRC) matching, which is devised to mitigate the adverse effects caused by noisy correspondence. Specifically, NeiRC first uses CLIP to compute similarity scores of image-text pairs and then partition these scores into clean and noisy distributions. NeiRC searches for similar semantic neighbors for each sample pair from relatively clean data. Subsequently, these neighbors are used by NeiRC to identify and replace noisy samples, thus helping the model obtain correct correspondences and maintain stable performance. Comprehensive experiments on the Flickr30K, MS-COCO, and Conceptual Captions datasets have demonstrated the effectiveness and robustness of NeiRC, and its superiority is particularly significant especially in high-noise conditions.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cross-Modal Matching with Noisy Correspondence via Neighbor Replacing

  • Ao Han,
  • Yang Cao

摘要

The performance of existing cross-modal matching techniques highly depends on the precise alignment of multimodal data. Nevertheless, in the process of data collection or annotation, certain data pairs may suffer from compromised correspondences, giving rise to the problem of noisy correspondence. To tackle this hurdle, we propose an innovative approach, namely Neighbor-Replacing Cross-modal (NeiRC) matching, which is devised to mitigate the adverse effects caused by noisy correspondence. Specifically, NeiRC first uses CLIP to compute similarity scores of image-text pairs and then partition these scores into clean and noisy distributions. NeiRC searches for similar semantic neighbors for each sample pair from relatively clean data. Subsequently, these neighbors are used by NeiRC to identify and replace noisy samples, thus helping the model obtain correct correspondences and maintain stable performance. Comprehensive experiments on the Flickr30K, MS-COCO, and Conceptual Captions datasets have demonstrated the effectiveness and robustness of NeiRC, and its superiority is particularly significant especially in high-noise conditions.