The field of vision-based Continuous Sign Language Recognition (CSLR) offers a promising solution to bridge the communication gap between the hearing and Deaf and Hard of Hearing (DHH) communities by enabling the translation of sign language videos into written or spoken language. Current state-of-the-art CSLR approaches often rely on complex deep learning architectures to capture the spatial and temporal features of sign language, but these complex models often come at a significant computational cost. Our work proposes an efficient framework that leverages the Temporal Shift Module (TSM) and Squeeze-and-Excitation (SE) module to enhance temporal modeling and feature representation. The TSM empowers 2D Convolutional Neural Networks (CNNs) to effectively capture temporal features without increasing computational resources, while the SE module adaptively selects informative features. Our proposed framework is designed to be both accurate and efficient. We evaluate our framework on the PHOENIX14 and PHOENIX14-T datasets, achieving competitive recognition performance compared to prior works while requiring only a small increase in computational resources.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Efficient Continuous Sign Language Recognition with Temporal Shift and Channel Attention

  • Nguyen Tu Nam,
  • Hiroki Takahashi

摘要

The field of vision-based Continuous Sign Language Recognition (CSLR) offers a promising solution to bridge the communication gap between the hearing and Deaf and Hard of Hearing (DHH) communities by enabling the translation of sign language videos into written or spoken language. Current state-of-the-art CSLR approaches often rely on complex deep learning architectures to capture the spatial and temporal features of sign language, but these complex models often come at a significant computational cost. Our work proposes an efficient framework that leverages the Temporal Shift Module (TSM) and Squeeze-and-Excitation (SE) module to enhance temporal modeling and feature representation. The TSM empowers 2D Convolutional Neural Networks (CNNs) to effectively capture temporal features without increasing computational resources, while the SE module adaptively selects informative features. Our proposed framework is designed to be both accurate and efficient. We evaluate our framework on the PHOENIX14 and PHOENIX14-T datasets, achieving competitive recognition performance compared to prior works while requiring only a small increase in computational resources.