Indian Sign Language Recognition at Word Level Using 3D-CNN
摘要
This study introduces a 3D convolutional neural network (3D-CNN) model for recognizing sign language expressions from video inputs, specifically tailored to assist communication for individuals with speech and hearing disabilities. The model employs spatiotemporal filters to extract both spatial and temporal features, processing detected faces in each frame to accurately classify sign expressions. We demonstrate our work on INCLUDE-50 and IRKSL datasets, obtain an accuracy of 94% on INCLUDE-50 dataset, and achieve 97% of accuracy on IRKSL dataset.