Historical states modeling for visual tracking
摘要
Extracting additional spatiotemporal information from video sequences is critical for accurately perceiving target appearance changes during visual tracking. However, most learning-based trackers utilize only a single search image and template from a video for training, resulting in a lack of temporal information and low data utilization. To address these issues, we present an innovative Trajectory Guided Tracking (TGTrack) framework, which leverages the historical states of the target to predict its current location. Specifically, we construct trajectory tokens derived from tracking results in historical frames, integrating the position and scale information of the target. We propose a trajectory prediction module to utilize these trajectory tokens to generate the potential scope of current target. Furthermore, to enhance the inference efficiency of the tracker, we eliminate manually customized heads and post-processing steps. Consequently, we achieve a good balance between inference speed and effectiveness. Extensive experimental results demonstrate that our TGTrack achieves leading performance across multiple benchmarks.