ViT Sign: An Effective Transformer-Based Approach for Sign Language Recognition
摘要
Sign language recognition (SLR) plays a critical role in enabling seamless communication for individuals with hearing impairments. This study introduces a novel SLR framework leveraging vision transformers (ViTs) as the primary deep learning architecture. ViTs, renowned for their efficacy in image classification, are employed to process video frames for robust sign recognition. The framework includes a preprocessing pipeline to normalize input data, enhancing the quality of the learned representations. The proposed methodology is evaluated on two benchmark datasets: the Malaysian sign language dataset and the Chinese sign language dataset. Experimental results demonstrate the model’s capability, achieving accuracy rates of 92.39% and 90.26% on the Chinese and Malaysian datasets, respectively. These results underscore the potential of ViT-based architectures in recognizing signs across diverse linguistic and cultural contexts, paving the way for advanced, inclusive SLR systems. Furthermore, this approach highlights the scalability and adaptability of transformers for multimodal gesture recognition tasks, setting a foundation for future research in sign language translation and real-time systems.