HCV: Lightweight Hybrid CNN-Vision Transformer for Visual Object Tracking
摘要
Visual object tracking is one of the most fundamental research in computer vision. Recent mainstream trackers prioritize accuracy, leading to issues such as prolonged computation time and substantial computational resources to achieve significant performance. To address this challenge, in this paper, we propose a lightweight single object tracking model named HCV. Through the proposed feature fusion module, we establish an interface between CNN and Vision Transformer (ViT), enabling the utilization of diverse hierarchical features from CNN while simultaneously acquiring global features through attention mechanisms. This approach maintains good tracking accuracy while reducing computational overhead and parameter requirements. We evaluate this idea on the UAV123, LaSOT, GOT-10k, and TrackingNet datasets, and demonstrate the efficiency of this lightweight tracking model.