Re-Identification Based on the Spatial-Temporal Fusion Network
摘要
Re-identification (ReID) in a large-scale camera network is critical in public safety, traffic control, and security. However, due to the ambiguous appearance of objects, the previous appearance-based ReID methods often fail to track objects across multiple cameras. To overcome this challenge, we propose a ReID based on a spatial-temporal fusion network that estimates a reliable camera network topology based on the adaptive Parzen window method and optimally combines the appearance and spatial-temporal similarities through a fusion network. The proposed methods demonstrated the best performance on the public vehicle dataset (VeRi776) with 99.7% rank-1 accuracy and on the person dataset (Market1501) with 99.11% rank-1 accuracy. The experimental results support that using spatial and temporal information for ReID can leverage the accuracy of appearance-based methods and effectively manage appearance ambiguities.