Lip reading decodes spoken language by observing lip movements, improving speech recognition in noisy environments. Current techniques rely primarily on visual features, which often result in suboptimal generalization to unseen speakers. Lip landmark features, which describe speaker-independent lip movements, can overcome this limitation. This paper presents a novel approach to constructing spatio-temporal features from landmarks. By using intra-frame Euclidean distances along with inter-frame velocities and accelerations, it effectively captures lip movements, thereby improving the generalization of lip-reading models to unseen speakers. The proposed multimodal fusion architecture integrates visual and landmark features through cross-attention mechanisms, following self-attention processing in each branch. Residual connections integrate pre- and post-fusion features to enhance comprehensive feature representation, and a feature enhancement strategy using random feature masking further augments the model’s generalization capabilities. A new dataset, CECC-VSR, is constructed, comprising a limited number of speakers recorded in real, complex indoor or outdoor environments, where speakers can freely move within the recording scenes. Experimental results on the GRID, LRW-ID, and CECC-VSR datasets indicate that the proposed method significantly improves generalization for unseen speakers. It reduces the WER by 4.16% on the GRID and increases accuracy by 1.97% on the LRW-ID, both compared to the baseline.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Advancing Lip-Reading for Unseen Speakers Through Fusion and Augmentation of Spatio-temporal Landmarks and Visual Features

  • Weihua Qiang,
  • Shizhan Chen,
  • Xingyu Zhang,
  • Yakun Zhang,
  • Xusheng Wang,
  • Changyan Zheng,
  • Liang Xie,
  • Ye Yan,
  • Erwei Yin

摘要

Lip reading decodes spoken language by observing lip movements, improving speech recognition in noisy environments. Current techniques rely primarily on visual features, which often result in suboptimal generalization to unseen speakers. Lip landmark features, which describe speaker-independent lip movements, can overcome this limitation. This paper presents a novel approach to constructing spatio-temporal features from landmarks. By using intra-frame Euclidean distances along with inter-frame velocities and accelerations, it effectively captures lip movements, thereby improving the generalization of lip-reading models to unseen speakers. The proposed multimodal fusion architecture integrates visual and landmark features through cross-attention mechanisms, following self-attention processing in each branch. Residual connections integrate pre- and post-fusion features to enhance comprehensive feature representation, and a feature enhancement strategy using random feature masking further augments the model’s generalization capabilities. A new dataset, CECC-VSR, is constructed, comprising a limited number of speakers recorded in real, complex indoor or outdoor environments, where speakers can freely move within the recording scenes. Experimental results on the GRID, LRW-ID, and CECC-VSR datasets indicate that the proposed method significantly improves generalization for unseen speakers. It reduces the WER by 4.16% on the GRID and increases accuracy by 1.97% on the LRW-ID, both compared to the baseline.