Attention to the Branches: A Comparative Analysis of FairMOT with Transformers on Fish Dataset
摘要
The application of Transformers in computer vision has gained momentum, even to the extent of revising the Vision Transformer (ViT) theory of abandoning CNNs, or to be exact CNN backbones for Transformer-based backbones. This research attempts to evaluate the efficiency backbones when incorporated into a re-ID-based model such as FairMOT which is traditionally trained using a CNN. We investigate how Transformer-based feature extraction impacts tracking performance, particularly for small and occluded objects such as fish in video data. Our findings indicate that while ViT backbones offer promising features, they do not yet surpass CNN-based methods in terms of tracking accuracy in regards to the FairMOT approach. This study highlights the need for further optimization of Transformer architectures.