TGTrack: Text Modality Autoregression and Generative Template Updating for Visual Object Tracking
摘要
In this paper, we introduce TGTrack, a novel visual tracking framework integrating Text Modality Autoregression and Generative Template Updating. TGTrack expands the latent feature space with an autoregressive decoder that models text modality information, encoding target trajectory coordinates and reconstructing update templates for temporal information. By converting trajectory coordinates into discrete tokens for a retention-based decoder, we enhance temporal modeling and awareness of the tracker. A novel generative template updating strategy is presented to handle challenges like object deformation and occlusion, reconstructing update templates directly instead of the previous discriminative approach. Experimental results demonstrate TGTrack’s competitiveness on benchmarks: achieving 72.2% \(SR_{0.75}\) on GOT-10k and 80.0% \(P_{Norm}\) on LaSOT, validating our framework’s effectiveness.