Spatial Temporal Signatures: A Hybrid CNN-LSTM Architecture for Improved Sign Language Recognition
摘要
Sign language is a type of communication in which communication is expressed through hand signs and gestures. The challenge of predicting sign language lies in the need for models to take into consideration the spatial complexity and temporal fluctuations inherent in gestural communication. Our research presents a novel approach to sign language translation that makes use of the complementary abilities of convolutional neural networks (CNNs) and long short-term memory (LSTM) networks. With our well thought-out hybrid architecture, temporal relationships are captured by LSTM layers and spatial data is extracted by CNN layers. By combining these two powerful components, we introduce a model that is able to understand and predict whole sign language gestures. Experimental results on multiple datasets show that our technology outperforms current approaches. In order to convert sign language precisely, it also emphasizes how important it is to combine spatial and temporal information. This research contributes to the state-of-the-art in sign language prediction and sheds light on a wider range of multimodal sequence-to-sequence problems.