Visual object tracking is one of the most fundamental research in computer vision. Recent mainstream trackers prioritize accuracy, leading to issues such as prolonged computation time and substantial computational resources to achieve significant performance. To address this challenge, in this paper, we propose a lightweight single object tracking model named HCV. Through the proposed feature fusion module, we establish an interface between CNN and Vision Transformer (ViT), enabling the utilization of diverse hierarchical features from CNN while simultaneously acquiring global features through attention mechanisms. This approach maintains good tracking accuracy while reducing computational overhead and parameter requirements. We evaluate this idea on the UAV123, LaSOT, GOT-10k, and TrackingNet datasets, and demonstrate the efficiency of this lightweight tracking model.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

HCV: Lightweight Hybrid CNN-Vision Transformer for Visual Object Tracking

  • Liang-Chia Chen,
  • Wei-Ta Chu

摘要

Visual object tracking is one of the most fundamental research in computer vision. Recent mainstream trackers prioritize accuracy, leading to issues such as prolonged computation time and substantial computational resources to achieve significant performance. To address this challenge, in this paper, we propose a lightweight single object tracking model named HCV. Through the proposed feature fusion module, we establish an interface between CNN and Vision Transformer (ViT), enabling the utilization of diverse hierarchical features from CNN while simultaneously acquiring global features through attention mechanisms. This approach maintains good tracking accuracy while reducing computational overhead and parameter requirements. We evaluate this idea on the UAV123, LaSOT, GOT-10k, and TrackingNet datasets, and demonstrate the efficiency of this lightweight tracking model.