<p>Sign Language Production (SLP) aims to generate a sign language sequence corresponding to a spoken language sequence. Using gloss sequences to generate pose sequences (G2P) is the core task of SLP. Current work ignores the joint constraints for precise finger angles, resulting in the formation of irrational finger positions. In this paper, we propose an angle and graph topology enhanced framework named AG-TEF. AG-TEF consists of a dual-channel mixed token processing unit that integrates a spatio-temporal transformer and a spatio-temporal graph convolutional network (GCN). The framework’s decoder is comprised of three principal components: an embedding layer, a Transformer-global, and a GCN-local. In the embedding layer, we introduce AngleGCN and JointGCN with GCN each with separate parameters. AngleGCN focuses on embedding angular information to correct finger angle deviations, while JointGCN handles joint position information to address occlusion issues. The remaining components, Transformer-global and GCN-local, combine to form a dual-channel mixed token processing unit that learns both global and local dependencies. In addition, attention fusion is applied to dual-channel and Transformer global features to create a fine representation of the fingers. The experimental results show that the BLEU-4 scores of AG-TEF on the PHOENIX14T and our self-collected Chinese vocabulary sign language dataset are 11.14 and 15.62.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Angle and graph topology enhanced framework with dual-channel mixed token progressing unit for sign language production

  • Yarun Yang,
  • Qingshan Wang,
  • Qi Wang,
  • Sheng Chen

摘要

Sign Language Production (SLP) aims to generate a sign language sequence corresponding to a spoken language sequence. Using gloss sequences to generate pose sequences (G2P) is the core task of SLP. Current work ignores the joint constraints for precise finger angles, resulting in the formation of irrational finger positions. In this paper, we propose an angle and graph topology enhanced framework named AG-TEF. AG-TEF consists of a dual-channel mixed token processing unit that integrates a spatio-temporal transformer and a spatio-temporal graph convolutional network (GCN). The framework’s decoder is comprised of three principal components: an embedding layer, a Transformer-global, and a GCN-local. In the embedding layer, we introduce AngleGCN and JointGCN with GCN each with separate parameters. AngleGCN focuses on embedding angular information to correct finger angle deviations, while JointGCN handles joint position information to address occlusion issues. The remaining components, Transformer-global and GCN-local, combine to form a dual-channel mixed token processing unit that learns both global and local dependencies. In addition, attention fusion is applied to dual-channel and Transformer global features to create a fine representation of the fingers. The experimental results show that the BLEU-4 scores of AG-TEF on the PHOENIX14T and our self-collected Chinese vocabulary sign language dataset are 11.14 and 15.62.