<p>Developing efficient and effective sign language recognition systems has drawn much interest as inclusive communication technologies grow in value. The subtle hand gestures, finger articulations, and dynamic stances of sign language hinder even achieving great accuracy and generalizability. This work develops the processing efficient and accuracy-boosting MobileNetv2 model to provide a deep learning-based solution to these difficulties. One can assess the proposed model’s performance, adaptability, and generalizability using top-tier models, ResNet-50, VGG-16, InceptionV3, EfficientNet, AlexNet, and Xception. The ArSL 2018 dataset is compared with these models. Furthermore, this study includes a cross-dataset evaluation conducted using the ArSL2L dataset. Furthermore, K-fold cross-valuation has been used to show the model’s generalizability; the model shows constant performance over all folds. With performance criteria ranging from 90 to 96%, results revealed that the model was far more resilient and versatile than rival models. The proposed model achieved a testing accuracy of 97.22%, demonstrating its effectiveness and robustness. To demonstrate the model’s real-time applicability, we evaluated latency and memory usage on a mobile device. The model achieved an average inference time of 42&#xa0;ms per frame with a lightweight footprint of 13.6&#xa0;MB, making it suitable for deployment on mobile platforms. Our study shows that the MobileNetv2 model performs well in identifying complicated gestures.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Advancing Arabic Sign Language Recognition: A Novel MobileNetv2-Based DL Framework with Superior Accuracy and Cross-Dataset Validation

  • Rabia Emhamed Al Mamlook,
  • Abeer Aljohani

摘要

Developing efficient and effective sign language recognition systems has drawn much interest as inclusive communication technologies grow in value. The subtle hand gestures, finger articulations, and dynamic stances of sign language hinder even achieving great accuracy and generalizability. This work develops the processing efficient and accuracy-boosting MobileNetv2 model to provide a deep learning-based solution to these difficulties. One can assess the proposed model’s performance, adaptability, and generalizability using top-tier models, ResNet-50, VGG-16, InceptionV3, EfficientNet, AlexNet, and Xception. The ArSL 2018 dataset is compared with these models. Furthermore, this study includes a cross-dataset evaluation conducted using the ArSL2L dataset. Furthermore, K-fold cross-valuation has been used to show the model’s generalizability; the model shows constant performance over all folds. With performance criteria ranging from 90 to 96%, results revealed that the model was far more resilient and versatile than rival models. The proposed model achieved a testing accuracy of 97.22%, demonstrating its effectiveness and robustness. To demonstrate the model’s real-time applicability, we evaluated latency and memory usage on a mobile device. The model achieved an average inference time of 42 ms per frame with a lightweight footprint of 13.6 MB, making it suitable for deployment on mobile platforms. Our study shows that the MobileNetv2 model performs well in identifying complicated gestures.