A novel key frame extraction using a deep learning model for sign language recognition on videos
摘要
Sign language recognition (SLR) plays a critical role in bridging communication gaps for the hearing- and speech-impaired community. However, current systems struggle with challenges such as redundant video frames, dynamic backgrounds, signer variations, and limited real-time accuracy. Motivated by the need for an efficient, scalable, and accurate solution, this study proposes a novel deep learning framework for robust keyframe extraction and SLR. The method begins with the preparation of two benchmark datasets, INCLUDE and INCLUDE-50, followed by action-based keyframe extraction using structural similarity (SSIM), cosine similarity, and MSE/MAE thresholds. A customized CNN model named SignKeyNet-AR is developed, integrating atrous convolutions, residual connections, and a ResNet-50 backbone to improve gesture classification. The use of IGMM for background subtraction enhances data quality, while Optical Flow aids normalization. Experimental results show superior performance across multiple metrics (accuracy: 99.05%, mAP: 97.92%), outperforming state-of-the-art KFE techniques. This approach provides an effective, real-time solution for accurate sign language video recognition.