Efficient Continuous Sign Language Recognition with Temporal Shift and Channel Attention
摘要
The field of vision-based Continuous Sign Language Recognition (CSLR) offers a promising solution to bridge the communication gap between the hearing and Deaf and Hard of Hearing (DHH) communities by enabling the translation of sign language videos into written or spoken language. Current state-of-the-art CSLR approaches often rely on complex deep learning architectures to capture the spatial and temporal features of sign language, but these complex models often come at a significant computational cost. Our work proposes an efficient framework that leverages the Temporal Shift Module (TSM) and Squeeze-and-Excitation (SE) module to enhance temporal modeling and feature representation. The TSM empowers 2D Convolutional Neural Networks (CNNs) to effectively capture temporal features without increasing computational resources, while the SE module adaptively selects informative features. Our proposed framework is designed to be both accurate and efficient. We evaluate our framework on the PHOENIX14 and PHOENIX14-T datasets, achieving competitive recognition performance compared to prior works while requiring only a small increase in computational resources.