Sign Language Recognition (SLR) is a crucial technological advancement that significantly enhances communication for Deaf and Hard of Hearing (DHH) individuals. Recognizing the need for robust and accurate systems in this field, this study introduces an innovative approach to SLR, focusing on a hybrid architecture that combines Convolutional Neural Networks (CNNs) and Long Short-Term Memory (LSTM) networks. This approach is specifically tailored to address the complexities of the Word-Level American Sign Language (WLASL) dataset. Our proposed model effectively captures the spatial and temporal nuances of sign language videos, achieving impressive performance metrics such as high accuracy, precision, recall, and F1-score. By leveraging the strengths of both CNNs and LSTMs, our model demonstrates a superior ability to recognize intricate gestures and actions, overcoming the limitations of traditional methods that often focus only on spatial or temporal aspects. The methodology integrates CNN features with LSTM temporal modeling, providing a robust framework for SLR. While the current model shows promising results, future improvements could include expanding the dataset and fine-tuning the architecture to further enhance performance. This study represents a significant advancement in SLR technology, with considerable potential for real-world applications in assistive technologies and accessibility. The implications of our findings extend beyond technological progress, offering greater inclusivity and accessibility for DHH individuals in various aspects of daily life.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comprehensive Sign Language Recognition Based on Hybrid Deep Learning Models: Integrating CNNs and RNNs/LSTMs for Enhanced Spatial-Temporal Analysis

  • Zizoune Aicha,
  • Zizoune Asmae,
  • Ziti Soumia,
  • Salah-Eddine Karima

摘要

Sign Language Recognition (SLR) is a crucial technological advancement that significantly enhances communication for Deaf and Hard of Hearing (DHH) individuals. Recognizing the need for robust and accurate systems in this field, this study introduces an innovative approach to SLR, focusing on a hybrid architecture that combines Convolutional Neural Networks (CNNs) and Long Short-Term Memory (LSTM) networks. This approach is specifically tailored to address the complexities of the Word-Level American Sign Language (WLASL) dataset. Our proposed model effectively captures the spatial and temporal nuances of sign language videos, achieving impressive performance metrics such as high accuracy, precision, recall, and F1-score. By leveraging the strengths of both CNNs and LSTMs, our model demonstrates a superior ability to recognize intricate gestures and actions, overcoming the limitations of traditional methods that often focus only on spatial or temporal aspects. The methodology integrates CNN features with LSTM temporal modeling, providing a robust framework for SLR. While the current model shows promising results, future improvements could include expanding the dataset and fine-tuning the architecture to further enhance performance. This study represents a significant advancement in SLR technology, with considerable potential for real-world applications in assistive technologies and accessibility. The implications of our findings extend beyond technological progress, offering greater inclusivity and accessibility for DHH individuals in various aspects of daily life.