<p>The objective of multi-object tracking (MOT) is to accurately detect and track all objects within a continuous sequence while maintaining unique identifiers for each object. Existing research predominantly relies on strong cues from motion and appearance to construct heuristic models, with limited attention given to weak cues arising from object overlap or shape changes as observed by advanced detectors. In this paper, we introduce a novel approach that leverages weak motion cues by adaptively integrating them into high-performance motion versus appearance-based methods. Moreover, by designing weak cue extraction and matching to run independently across targets, our method inherently supports parallelism and GPU acceleration, enabling efficient high-resolution video tracking and showing strong HPC potential. Building upon the appearance-based motion method Deep OC-SORT (in: IEEE International Conference on Image Processing, IEEE, 2023), our approach achieves superior performance on the challenging DanceTrack (in: Proceedings of the IEEE/CVF Conference on Computer 367 Vision and Pattern Recognition, 2022) benchmark, attaining a HOTA score of 61.9. Furthermore, compared to more complex methods, our approach achieves HOTA scores of 65.4 and 64.3 on the MOT17 (Milan in arXiv preprint , 2016) and MOT20 (Dendorfer in arXiv preprint, 2020) benchmarks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-cue SORT: integrating weak cues with appearance and motion for multi-object tracking

  • Hong Liang,
  • Mingchen Xu,
  • Qian Zhang,
  • Mingwen Shao

摘要

The objective of multi-object tracking (MOT) is to accurately detect and track all objects within a continuous sequence while maintaining unique identifiers for each object. Existing research predominantly relies on strong cues from motion and appearance to construct heuristic models, with limited attention given to weak cues arising from object overlap or shape changes as observed by advanced detectors. In this paper, we introduce a novel approach that leverages weak motion cues by adaptively integrating them into high-performance motion versus appearance-based methods. Moreover, by designing weak cue extraction and matching to run independently across targets, our method inherently supports parallelism and GPU acceleration, enabling efficient high-resolution video tracking and showing strong HPC potential. Building upon the appearance-based motion method Deep OC-SORT (in: IEEE International Conference on Image Processing, IEEE, 2023), our approach achieves superior performance on the challenging DanceTrack (in: Proceedings of the IEEE/CVF Conference on Computer 367 Vision and Pattern Recognition, 2022) benchmark, attaining a HOTA score of 61.9. Furthermore, compared to more complex methods, our approach achieves HOTA scores of 65.4 and 64.3 on the MOT17 (Milan in arXiv preprint , 2016) and MOT20 (Dendorfer in arXiv preprint, 2020) benchmarks.