Multi-vehicle tracking in highway roadside surveillance videos via spatial–temporal fusion network with adaptive decay memory
摘要
This paper presents spatial–temporal fusion network with adaptive decay memory for multi-vehicle tracking in highway scenarios (STFN-ADM-MTHS), a spatial–temporal fusion network with adaptive decay memory for robust multi-vehicle tracking in highway surveillance scenarios. Addressing challenges posed by camera vibrations, and occlusions, our framework includes four parts. Firstly, a last frame newborn module incorporating detection query self-attention to enhance detection-track coordination through temporal context; secondly, repulsion loss for improved occlusion handling via bounding box repulsion constraints; thirdly, spatial-guided cross-temporal feature aggregation (SCFA) using inter-frame interpolation and attention mechanisms to learn vibration-induced motion patterns; and lastly, adaptive decay memory (ADM) that dynamically adjusts memory retention based on motion stability, prioritizing recent observations during stable periods while relying on historical data during occlusions. We contribute the MTHS-SEU dataset containing 21,558 highway surveillance images with diverse environmental challenges. Comprehensive evaluations across MTHS-SEU, UA-DETRAC, and KITTI benchmarks demonstrate superior performance, achieving 72.197% HOTA on our dataset. Ablation studies confirm individual module effectiveness, with SCFA and ADM particularly effective for vibration compensation and occlusion handling, respectively. Given the high-dimensional spatiotemporal features and the need for long-sequence trajectory modeling under vibration and occlusion, our framework leverages GPU-accelerated parallel processing to sustain real-time inference. This capability is critical for meeting the safety-critical response needs of intelligent transportation systems. The framework maintains real-time capability at 10.5 FPS, showing significant improvements over state-of-the-art trackers in association metrics while balancing detection accuracy through joint query decoding optimization. The results show that the STFN-ADM-MTHS model effectively addresses the inter-frame trajectory displacement caused by camera vibrations.