<p>Text-based person search (TBPS) faces challenges in aligning features of both text and image modalities, with pervious methods assuming correct alignment of training image-text pairs. However, issues such as image blurriness and annotation errors lead to noise correspondence problems, resulting in low correlation and pseudo-correlation in image-text pairs. To address these under strong noise interference and learn robust visual-semantic associations, we propose a dual-evidence-based cross-modal mask alignment (DEMA) framework, which includes a consensus evidence division (CED) module and a cross-modal masking (CMM) module. The CED filters out correct and reliable visual-semantic associations through a dual-granularity embedding evidence learning network. Meanwhile, the CMM utilizes bidirectional masking and "error word" correction to focus on masked text and images, reducing noise interference and strengthening the latent association between image and text. Extensive experiments on the CUHK-PEDES, ICFG-PEDES, and RSTPReid datasets demonstrate the effectiveness of our DEMA method.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Noise correspondence with evidence learning for text-based person search

  • Yihan Xie,
  • Baohua Zhang,
  • Yang Li,
  • Chongrui Shan,
  • Shun Wang,
  • Jiale Zhang

摘要

Text-based person search (TBPS) faces challenges in aligning features of both text and image modalities, with pervious methods assuming correct alignment of training image-text pairs. However, issues such as image blurriness and annotation errors lead to noise correspondence problems, resulting in low correlation and pseudo-correlation in image-text pairs. To address these under strong noise interference and learn robust visual-semantic associations, we propose a dual-evidence-based cross-modal mask alignment (DEMA) framework, which includes a consensus evidence division (CED) module and a cross-modal masking (CMM) module. The CED filters out correct and reliable visual-semantic associations through a dual-granularity embedding evidence learning network. Meanwhile, the CMM utilizes bidirectional masking and "error word" correction to focus on masked text and images, reducing noise interference and strengthening the latent association between image and text. Extensive experiments on the CUHK-PEDES, ICFG-PEDES, and RSTPReid datasets demonstrate the effectiveness of our DEMA method.