The field of RGBD object tracking is gaining increasing attention due to the ability of depth sensors to capture spatial information. Consequently, the inclusion of depth modality enables trackers to perform better in scenarios with occlusion, complex backgrounds, and dark environments. However, how to fully exploit both RGB and depth modality is still a challenging problem, since the features encoded in these two modalities exhibit both commonalities and individualities. Some researchers integrate the two modalities using simple methods without evaluating the effectiveness of these operations. Another important issue which is usually overlooked by exiting literatures is the impact of imaging condition such as illumination, since different modalities exhibit varying degrees of effectiveness under different imaging conditions. To address these issues, we first conducted a thorough study of existing fusion operators and found that most of them can be attributed to a type of \(l^p\) -norm. Subsequently, we propose an \(l^2\) -norm based fusion method, which can balance the description of the individual characteristics of modalities and the common features between modalities during the fusion process. Furthermore, considering the varying effectiveness of different modalities under different lighting conditions, we introduce the illumination counter to guide the tracker. Both of the \(l^2\) -norm fusion operator and the illumination counter are parameter-free yet bring about significant improvements on tracking performance on publicly available RGB-D tracking datasets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

L2FIG-Tracker: L2-Norm Based Fusion with Illumination Guidance for RGB-D Object Tracking

  • Jintao Su,
  • Ye Liu,
  • Shitao Song

摘要

The field of RGBD object tracking is gaining increasing attention due to the ability of depth sensors to capture spatial information. Consequently, the inclusion of depth modality enables trackers to perform better in scenarios with occlusion, complex backgrounds, and dark environments. However, how to fully exploit both RGB and depth modality is still a challenging problem, since the features encoded in these two modalities exhibit both commonalities and individualities. Some researchers integrate the two modalities using simple methods without evaluating the effectiveness of these operations. Another important issue which is usually overlooked by exiting literatures is the impact of imaging condition such as illumination, since different modalities exhibit varying degrees of effectiveness under different imaging conditions. To address these issues, we first conducted a thorough study of existing fusion operators and found that most of them can be attributed to a type of \(l^p\) -norm. Subsequently, we propose an \(l^2\) -norm based fusion method, which can balance the description of the individual characteristics of modalities and the common features between modalities during the fusion process. Furthermore, considering the varying effectiveness of different modalities under different lighting conditions, we introduce the illumination counter to guide the tracker. Both of the \(l^2\) -norm fusion operator and the illumination counter are parameter-free yet bring about significant improvements on tracking performance on publicly available RGB-D tracking datasets.