<p>The proposed work addresses Continuous Sign Language Recognition and Translation (CSLRT) by introducing the Three-Stream Multimodal Occlusion-Resilient Transformer (T-MORT) model. T-MORT leverages three parallel encoder transformers to extract complementary features from visual, gesture, and emotion cues in sign language videos. These multi-modal representations are fused using a late fusion strategy, which preserves modality-specific information and enhances robustness to occlusions. A dual-decoder framework is employed: one decoder autoregressively generates gloss sequences to provide linguistic priors, while the second produces the final spoken language translation. To enhance efficiency, a dynamic keypoint extraction-based frame sampling algorithm, Spatio-Temporal Keypoint-Guided Sampling (STKGS), adaptively selects frames, reducing computational overhead. Extensive experiments on the PHOENIX14T dataset demonstrate state-of-the-art performance, surpassing existing CSLRT benchmarks. The proposed T-MORT model yields a BLEU-4 score of 25.07 and a ROUGE-L of 49.12.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhanced sign language translation using three stream multimodal occlusion resilient transformer

  • Rina Damdoo,
  • Praveen Kumar

摘要

The proposed work addresses Continuous Sign Language Recognition and Translation (CSLRT) by introducing the Three-Stream Multimodal Occlusion-Resilient Transformer (T-MORT) model. T-MORT leverages three parallel encoder transformers to extract complementary features from visual, gesture, and emotion cues in sign language videos. These multi-modal representations are fused using a late fusion strategy, which preserves modality-specific information and enhances robustness to occlusions. A dual-decoder framework is employed: one decoder autoregressively generates gloss sequences to provide linguistic priors, while the second produces the final spoken language translation. To enhance efficiency, a dynamic keypoint extraction-based frame sampling algorithm, Spatio-Temporal Keypoint-Guided Sampling (STKGS), adaptively selects frames, reducing computational overhead. Extensive experiments on the PHOENIX14T dataset demonstrate state-of-the-art performance, surpassing existing CSLRT benchmarks. The proposed T-MORT model yields a BLEU-4 score of 25.07 and a ROUGE-L of 49.12.