Noise correspondence with evidence learning for text-based person search
摘要
Text-based person search (TBPS) faces challenges in aligning features of both text and image modalities, with pervious methods assuming correct alignment of training image-text pairs. However, issues such as image blurriness and annotation errors lead to noise correspondence problems, resulting in low correlation and pseudo-correlation in image-text pairs. To address these under strong noise interference and learn robust visual-semantic associations, we propose a dual-evidence-based cross-modal mask alignment (DEMA) framework, which includes a consensus evidence division (CED) module and a cross-modal masking (CMM) module. The CED filters out correct and reliable visual-semantic associations through a dual-granularity embedding evidence learning network. Meanwhile, the CMM utilizes bidirectional masking and "error word" correction to focus on masked text and images, reducing noise interference and strengthening the latent association between image and text. Extensive experiments on the CUHK-PEDES, ICFG-PEDES, and RSTPReid datasets demonstrate the effectiveness of our DEMA method.