Enhancing Communication with Advanced CNN Models for Recognizing American Sign Language
摘要
The research work endeavors to enhance automatic sign language interpretation and communication by leveraging Convolutional Neural Networks (CNN) for the recognition and classification of static hand gestures representing the 26 English alphabets (A-Z) and limited to 10 numbers (0–9). First, a comprehensive dataset of hand gestures for the English alphabet and numbers is created. Next, images are carefully pre-processed by cleaning and resizing them. Finally, models are trained using various CNN architectures, including LeNet-5, MobileNetV2, and a custom network. To identify and categorise hand gestures, the dataset is thoroughly trained. The models’ accuracy and performance are evaluated on additional datasets. Approximately 96.70% of the tested dataset is recognised accurately by the trained CNN models. Using an ensemble approach and efficient picture pre-processing techniques are important variables that contribute to the strength and durability of the models. Confusion matrix analysis, which demonstrates considerable differences between accurate and inaccurate identifications, highlights the models’ efficacy. An effective method for translating static hand motions into text for sign language was created by the study team. Right now, the system can recognise letters and numbers with accuracy from hand gestures, but it cannot understand dynamic movements or full-body language. It is limited to static motions. The system's vocabulary might be increased and its capacity to identify dynamic movements could be strengthened in future iterations. With this advancement, the deaf and hard-of-hearing community's communication accessibility could be greatly enhanced by sophisticated technologies that comprehend and facilitate sign language communication better.