Pose-invariant gait recognition using an enhanced inception-ResNet and vision transformer framework
摘要
Gait recognition is an emerging biometric technique with wide-ranging applications in surveillance, security, and healthcare. Despite its advantages over traditional biometric modalities, performance remains sensitive to factors such as viewpoint variations, occlusions, and posture changes. To address these challenges, this study proposes a hybrid gait recognition framework that integrates Inception-ResNet and Vision Transformer (ViT) architectures for robust, pose-invariant feature extraction. The framework employs OpenPose to estimate skeletal keypoints, which are processed through a dual-path network: the Inception-ResNet branch captures multi-scale spatial representations of local gait patterns, while the ViT branch models long-range temporal dependencies via self-attention. This combination enables a comprehensive representation of both fine-grained spatial features and global temporal dynamics. Experiments conducted on benchmark datasets—CASIA-B, OU-MVLP, GREW and Gait3D—demonstrate that the proposed method surpasses state-of-the-art approaches, including GaitSet, GaitPart, and GaitGL, particularly under cross-view and occluded conditions. The framework achieves Rank-1 accuracy of 98.2% on CASIA-B under normal walking, 92.7% on OU-MVLP, 85.4% on GREW and 79.4% on Gait3D highlighting its robustness and practical potential for real-world applications.