Language-Assisted Cross-Modal Single Object Tracking
摘要
Natural language-assisted multi-modal object tracking can theoretically help reposition targets in the above scenarios by adding text description information. However, single object tracking datasets generally lack text descriptions, which reduces the scope of multi-modal tracking. Due to the transformation of the tracking target and background during the tracking process, it is challenging to maintain semantic consistency between text descriptions and images. In order to solve the problems of multi-modal tracking, this paper proposes a multi-modal single object tracking algorithm. This algorithm uses the image caption module to automatically generate text descriptions, which solves the problem of missing text descriptions of the initial images of some datasets. The algorithm uses a layer-by-layer fusion information module that uses visual features to pay different degrees of attention to text features in different layers of the visual information extraction and fusion module. The visual feature information is extracted and fused in the process of image-to-visual information interaction. Experimental results on single object tracking datasets show that it has achieved good results.