Natural language-assisted multi-modal object tracking can theoretically help reposition targets in the above scenarios by adding text description information. However, single object tracking datasets generally lack text descriptions, which reduces the scope of multi-modal tracking. Due to the transformation of the tracking target and background during the tracking process, it is challenging to maintain semantic consistency between text descriptions and images. In order to solve the problems of multi-modal tracking, this paper proposes a multi-modal single object tracking algorithm. This algorithm uses the image caption module to automatically generate text descriptions, which solves the problem of missing text descriptions of the initial images of some datasets. The algorithm uses a layer-by-layer fusion information module that uses visual features to pay different degrees of attention to text features in different layers of the visual information extraction and fusion module. The visual feature information is extracted and fused in the process of image-to-visual information interaction. Experimental results on single object tracking datasets show that it has achieved good results.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Language-Assisted Cross-Modal Single Object Tracking

  • Yongkun Yang,
  • Jinhai Xiang,
  • Zhongmin Chen

摘要

Natural language-assisted multi-modal object tracking can theoretically help reposition targets in the above scenarios by adding text description information. However, single object tracking datasets generally lack text descriptions, which reduces the scope of multi-modal tracking. Due to the transformation of the tracking target and background during the tracking process, it is challenging to maintain semantic consistency between text descriptions and images. In order to solve the problems of multi-modal tracking, this paper proposes a multi-modal single object tracking algorithm. This algorithm uses the image caption module to automatically generate text descriptions, which solves the problem of missing text descriptions of the initial images of some datasets. The algorithm uses a layer-by-layer fusion information module that uses visual features to pay different degrees of attention to text features in different layers of the visual information extraction and fusion module. The visual feature information is extracted and fused in the process of image-to-visual information interaction. Experimental results on single object tracking datasets show that it has achieved good results.