Dual-Direction Spatio-Temporal Perception Network for Video Person Re-Identification
摘要
The mechanism for extracting effective information from video sequences has consistently been a key technology for video person re-identification. Traditional methods employing pooling fusion only capture features representing a global scale, resulting in insufficient feature representation and weakened discriminative ability. This paper proposes a Dual-direction Spatio-temporal Perception Network (DSTPN) to enhance the feature representation. Specifically, the proposed method incorporates the global feature obtained by pooling methods and richer spatio-temporal information extracted by the forward and backward branches, thereby augmenting the discriminative capacity of the features. Experiments on public benchmarks and ablation studies illustrate the efficacy of the method in this paper.