The application of Transformers in computer vision has gained momentum, even to the extent of revising the Vision Transformer (ViT) theory of abandoning CNNs, or to be exact CNN backbones for Transformer-based backbones. This research attempts to evaluate the efficiency backbones when incorporated into a re-ID-based model such as FairMOT which is traditionally trained using a CNN. We investigate how Transformer-based feature extraction impacts tracking performance, particularly for small and occluded objects such as fish in video data. Our findings indicate that while ViT backbones offer promising features, they do not yet surpass CNN-based methods in terms of tracking accuracy in regards to the FairMOT approach. This study highlights the need for further optimization of Transformer architectures.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Attention to the Branches: A Comparative Analysis of FairMOT with Transformers on Fish Dataset

  • Karim Anwar,
  • Seyed Sahand Mohammadi Ziabari

摘要

The application of Transformers in computer vision has gained momentum, even to the extent of revising the Vision Transformer (ViT) theory of abandoning CNNs, or to be exact CNN backbones for Transformer-based backbones. This research attempts to evaluate the efficiency backbones when incorporated into a re-ID-based model such as FairMOT which is traditionally trained using a CNN. We investigate how Transformer-based feature extraction impacts tracking performance, particularly for small and occluded objects such as fish in video data. Our findings indicate that while ViT backbones offer promising features, they do not yet surpass CNN-based methods in terms of tracking accuracy in regards to the FairMOT approach. This study highlights the need for further optimization of Transformer architectures.