<p>Vision Transformers have successfully captured long-range dependencies across various computer vision tasks, including Person Re-Identification (PReID). This paper proposed SwinReID, a novel PReID framework based on the Swin Transformer (ST) architecture. Unlike traditional convolutional neural networks (CNNs) that process entire images, SwinReID divides images into small, non-overlapping patches. This patch-based strategy enables a comprehensive understanding of global image context while preserving fine-grained spatial details, addressing challenges such as occlusions and appearance variations in PReID tasks. Conventional CNNs struggle with long-range dependencies due to their localized receptive fields and downsampling operations, limiting their effectiveness in complex scenarios. SwinReID overcomes these limitations by integrating an advanced transformer-based backbone and a lightweight channel attention (CA) module, which selectively improves discriminative feature channels. This refinement leads to better feature representation and robustness, particularly in challenging environments. SwinReID achieves state-of-the-art performance in benchmark datasets, with rank-1 accuracies of 95.5% in Market1501, 89.9% in DukeMTMC-reID, 64.1% in Occluded-Duke, and 91.63% in MSMT17, demonstrating its effectiveness in various and challenging PReID scenarios.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Swin transformer with attention mechanism: a novel framework for person re-identification

  • Tariq Ali Arain,
  • Pengcheng Zhang,
  • Qing Meng,
  • Abdullahi Uwaisu Muhammad

摘要

Vision Transformers have successfully captured long-range dependencies across various computer vision tasks, including Person Re-Identification (PReID). This paper proposed SwinReID, a novel PReID framework based on the Swin Transformer (ST) architecture. Unlike traditional convolutional neural networks (CNNs) that process entire images, SwinReID divides images into small, non-overlapping patches. This patch-based strategy enables a comprehensive understanding of global image context while preserving fine-grained spatial details, addressing challenges such as occlusions and appearance variations in PReID tasks. Conventional CNNs struggle with long-range dependencies due to their localized receptive fields and downsampling operations, limiting their effectiveness in complex scenarios. SwinReID overcomes these limitations by integrating an advanced transformer-based backbone and a lightweight channel attention (CA) module, which selectively improves discriminative feature channels. This refinement leads to better feature representation and robustness, particularly in challenging environments. SwinReID achieves state-of-the-art performance in benchmark datasets, with rank-1 accuracies of 95.5% in Market1501, 89.9% in DukeMTMC-reID, 64.1% in Occluded-Duke, and 91.63% in MSMT17, demonstrating its effectiveness in various and challenging PReID scenarios.