<p>Text-to-Image Person Re-identification (TI-ReID) has become a significant research field that focuses on retrieving pedestrian images from a gallery based on textual descriptions. Despite considerable progress, effectively aligning cross-modal features remains a fundamental challenge. Existing approaches often assume perfectly matched image-text pairs during training, yet such ideal conditions are difficult to achieve in practice, where visual ambiguities and annotation noise inevitably introduce noisy correspondences that degrade model performance. To address these challenges and learn robust visual-semantic associations under noisy conditions, we propose an Uncertainty-Guided Collaborative Learning (UGCL) framework. The framework consists of three main components. The Precise Correspondence Identification (PCI) module partitions the data into clean, uncertain, and noisy subsets through multi-level sample division. The Uncertainty-Guided Alignment (UGA) module models sample uncertainty and adaptively assigns weights to strengthen hard positive alignment while suppressing mismatched negatives. The Dynamic Pseudo-Label Rectification (DPR) module progressively updates pseudo-labels for noisy samples to enhance model robustness. Experimental results on multiple benchmark datasets demonstrate that the UGCL framework exhibits favorable robustness and competitive performance under both noisy and noise-free correspondence conditions.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Uncertainty-guided collaborative learning for noise-robust text-to-image person re-identification

  • Wen Ting,
  • Shidu Dong,
  • Yuzhi Zhang,
  • Zhenfang Yuan

摘要

Text-to-Image Person Re-identification (TI-ReID) has become a significant research field that focuses on retrieving pedestrian images from a gallery based on textual descriptions. Despite considerable progress, effectively aligning cross-modal features remains a fundamental challenge. Existing approaches often assume perfectly matched image-text pairs during training, yet such ideal conditions are difficult to achieve in practice, where visual ambiguities and annotation noise inevitably introduce noisy correspondences that degrade model performance. To address these challenges and learn robust visual-semantic associations under noisy conditions, we propose an Uncertainty-Guided Collaborative Learning (UGCL) framework. The framework consists of three main components. The Precise Correspondence Identification (PCI) module partitions the data into clean, uncertain, and noisy subsets through multi-level sample division. The Uncertainty-Guided Alignment (UGA) module models sample uncertainty and adaptively assigns weights to strengthen hard positive alignment while suppressing mismatched negatives. The Dynamic Pseudo-Label Rectification (DPR) module progressively updates pseudo-labels for noisy samples to enhance model robustness. Experimental results on multiple benchmark datasets demonstrate that the UGCL framework exhibits favorable robustness and competitive performance under both noisy and noise-free correspondence conditions.