Adjacent memory segmentation networks for robust visual tracking
摘要
Current high-performance trackers typically follow an offline matching-based paradigm, where the target states in subsequent frames are inferred based on the target template from the initial frame, which makes it difficult for the tracker to adapt to target appearance changes and discriminate distractors on the fly. Moreover, they only learn to associate the tracked targets across temporal frames, but ignore to exploit the space-time correspondence of other contents. To alleviate above issues, we propose a novel tracking framework based on adjacent memory segmentation networks, termed as TAMS, which provides up-to-date target appearance information as well as the background contents around the target. Specifically, TAMS leverages the previous frame with the estimated target state to prompt the accuracy of current target estimation, enabling the tracker to adapt to the target appearance change well. Moreover, we introduce a self-supervised correspondence learning method to explore the intrinsic coherence in videos, which allows background pixels/patches in the current frame have the chance to find true correspondences in previous frames, reducing the cases of incorrect matching with the tracked target. Finally, in order to achieve an accurate description of the target, TAMS produces an accurate segmentation mask and bounding box jointly, that clearly distinguishing the target from background content. The extensive experimental results on seven mainstream visual tracking benchmarks show that the proposed tracker achieves promising tracking performance.