A Lightweight Multi-modal Dynamic Fusion Network Based Method for Person Re-identification in Videos
摘要
Aiming at the image quality degradation and occlusion problems in video pedestrian re-identification, this paper proposes a lightweight multimodal feature learning and fusion model (EMAFNet). Firstly, this paper introduces the Contrast Constrained Adaptive Histogram Equalization (CLAHE) image enhancement technique, which effectively improves the details in areas of degraded image quality by enhancing the local contrast of the image. Then, through the joint modeling of visual appearance, gait and optical flow modalities, the complementarity and semantic relevance of multimodal information are fully exploited. We design a feature alignment module based on dynamic weight adjustment and cross-modal alignment is designed to realize fine-grained inter-modal information interaction and effectively enhance the robustness to complex scenes. Finally, this paper introduces an improved fast cross-attention mechanism to capture key correlation features between different modalities. Meanwhile, an effective occlusion region completion strategy is proposed to complete the missing regions due to occlusion by utilizing the time series context information. The experimental results show that the model in this paper achieves Rank-1 accuracies of 94.9%, 94.7% and 78.3% on the MARS, iLIDS-VID and Occluded-Duke datasets, respectively. Meanwhile, the inference speed reaches 40 FPS and the number of parameters is reduced to 36 M. It proves the effectiveness of the proposed method in complex scenarios, especially in image quality degradation and occlusion problems.