Transformer-Based Multimodal Spatial-Temporal Fusion for Gait Recognition
摘要
Gait recognition is a technique for identifying individuals by analyzing the motion characteristics of the human body while walking. Model-based and appearance-based methods are the two mainstream gait recognition methods, to combine the advantages of both, we fuse silhouette and posture features to propose a multimodal gait recognition method. Considering the powerful cross-attention mechanism and aggregation capability of the Transformer, we input the posture features containing coordinate and joint information and the fine-grained silhouette-based feature representation into the Transformer structure to realize the effective fusion of multimodal spatial features. Considering that features of different time scales will have an impact on the final gait recognition effect, we perform feature fusion of short-term spatial-temporal features and long-timescale spatial features to enhance the robustness of gait recognition. The final experimental results on a public dataset demonstrate the effectiveness of the proposed method. Compared with the traditional single-modal approach, the framework can significantly improve the performance and accuracy of gait recognition, providing a novel and effective solution for practical gait recognition applications.