<p>Automatic sign language generation has become an important assistive technology for bridging the communication gap between deaf or hard-of-hearing communities and the hearing majority, and for enabling accessible information services such as news broadcasting and online education. However, most existing approaches either rely heavily on gloss-level intermediate annotations or, when discarding glosses, still struggle with precise semantic–action alignment and with severe error accumulation in autoregressive decoding, which together limit the naturalness, scalability, and semantic fidelity of generated sign language videos. This paper proposes the Cross-modal Contrastive Creator (CCC), an end-to-end framework for fully gloss-free text-to-sign generation designed for generating semantically consistent sign language videos directly from text. Addressing key challenges in existing methods, such as semantic–action misalignment and error accumulation in autoregressive models, the CCC model incorporates a three-stage approach: first, an action reconstruction module based on MotionVQVAE that compresses continuous pose sequences into a discrete motion codebook while preserving fine-grained spatial–temporal dynamics; second, a multimodal fusion mechanism that couples a pre-trained text encoder with a false-negative-aware cross-modal contrastive objective and a parallel motion reconstruction task, thereby enforcing robust alignment between text and motion features; and third, an autoregressive generator built on a Transformer architecture that reuses the learned motion codebook as its embedding space and adopts a masking-based training strategy to effectively mitigate error accumulation during decoding. Furthermore, we introduce a normalized evaluation protocol that rescales BLEU and ROUGE scores with respect to a gloss-based reference system, enabling fairer comparison across different text-to-sign generation models. Experimental results on the RWTH-PHOENIX-Weather 2014T dataset demonstrate that CCC outperforms state-of-the-art methods in sign language generation. Specifically, our method achieves absolute scores of 39.43 for BLEU-1, 25.21 for BLEU-4, and 49.76 for ROUGE-L on the test set, representing improvements of 9.84%, 89.56%, and 35.35% over the previous best gloss-free method T2S-GPT, which obtained scores of 33.16, 11.87, and 34.65 respectively. Under the proposed normalized evaluation protocol, CCC achieves 90.06% for BLEU-1, 84.30% for BLEU-4, and 90.10% for ROUGE-L relative to the gloss-based baseline, outperforming T2S-GPT by 2.22, 1.52, and 1.64 percentage points respectively. Ablation studies confirm the effectiveness of each module, and comparison with gloss-dependent approaches shows that CCC provides more natural and accurate sign language generation while completely removing the need for intermediate gloss annotations.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CCC: cross-modal contrastive creator for end-to-end sign language generation

  • Wang Yi,
  • Ying Zhang,
  • Lu Meng,
  • Chengchen Cao,
  • Xuejie Lin,
  • Shuoqian Gao

摘要

Automatic sign language generation has become an important assistive technology for bridging the communication gap between deaf or hard-of-hearing communities and the hearing majority, and for enabling accessible information services such as news broadcasting and online education. However, most existing approaches either rely heavily on gloss-level intermediate annotations or, when discarding glosses, still struggle with precise semantic–action alignment and with severe error accumulation in autoregressive decoding, which together limit the naturalness, scalability, and semantic fidelity of generated sign language videos. This paper proposes the Cross-modal Contrastive Creator (CCC), an end-to-end framework for fully gloss-free text-to-sign generation designed for generating semantically consistent sign language videos directly from text. Addressing key challenges in existing methods, such as semantic–action misalignment and error accumulation in autoregressive models, the CCC model incorporates a three-stage approach: first, an action reconstruction module based on MotionVQVAE that compresses continuous pose sequences into a discrete motion codebook while preserving fine-grained spatial–temporal dynamics; second, a multimodal fusion mechanism that couples a pre-trained text encoder with a false-negative-aware cross-modal contrastive objective and a parallel motion reconstruction task, thereby enforcing robust alignment between text and motion features; and third, an autoregressive generator built on a Transformer architecture that reuses the learned motion codebook as its embedding space and adopts a masking-based training strategy to effectively mitigate error accumulation during decoding. Furthermore, we introduce a normalized evaluation protocol that rescales BLEU and ROUGE scores with respect to a gloss-based reference system, enabling fairer comparison across different text-to-sign generation models. Experimental results on the RWTH-PHOENIX-Weather 2014T dataset demonstrate that CCC outperforms state-of-the-art methods in sign language generation. Specifically, our method achieves absolute scores of 39.43 for BLEU-1, 25.21 for BLEU-4, and 49.76 for ROUGE-L on the test set, representing improvements of 9.84%, 89.56%, and 35.35% over the previous best gloss-free method T2S-GPT, which obtained scores of 33.16, 11.87, and 34.65 respectively. Under the proposed normalized evaluation protocol, CCC achieves 90.06% for BLEU-1, 84.30% for BLEU-4, and 90.10% for ROUGE-L relative to the gloss-based baseline, outperforming T2S-GPT by 2.22, 1.52, and 1.64 percentage points respectively. Ablation studies confirm the effectiveness of each module, and comparison with gloss-dependent approaches shows that CCC provides more natural and accurate sign language generation while completely removing the need for intermediate gloss annotations.