Audio Visual Speech Recognition (AVSR) has witnessed significant advancements by deploying sophisticated neural networks, particularly Convolutional Neural Networks (CNNs) and Transformers, to improve automatic speech recognition, lip-reading, and audio-visual scene understanding. However, challenges persist in effectively integrating Transformers—known for their ability to capture global context—into AVSR systems due to their substantial computational demands. This paper introduces a novel approach utilizing a modified version of the Conformer architecture, termed the Taylor-Expanded Linearized Conformer, designed to address these challenges. The proposed method retains the capability to model long-range dependencies while ensuring computational complexity remains linear. Additionally, a new architecture is introduced to mitigate feature collapse in Conformers by enhancing the nonlinearity in both the Multi-Head Attention and Feed-Forward Network components. The proposed architecture demonstrates superior efficiency and effectiveness, achieving state-of-the-art performance on the Lip Reading Sentences 2 (LRS2) and Lip Reading Sentences 3 (LRS3) datasets. An ablation study further validates the contributions of each component to the overall performance. The proposed model, which integrates Taylor-expanded attention and enhanced modules, exhibits robustness to noise and outperforms existing methods, offering a promising direction for the advancement of AVSR systems.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Conformer-Based Audio Visual Speech Recognition with Taylor Attention

  • Yewei Xiao,
  • Jian Huang,
  • Xuanming Liu,
  • Aosu Zhu

摘要

Audio Visual Speech Recognition (AVSR) has witnessed significant advancements by deploying sophisticated neural networks, particularly Convolutional Neural Networks (CNNs) and Transformers, to improve automatic speech recognition, lip-reading, and audio-visual scene understanding. However, challenges persist in effectively integrating Transformers—known for their ability to capture global context—into AVSR systems due to their substantial computational demands. This paper introduces a novel approach utilizing a modified version of the Conformer architecture, termed the Taylor-Expanded Linearized Conformer, designed to address these challenges. The proposed method retains the capability to model long-range dependencies while ensuring computational complexity remains linear. Additionally, a new architecture is introduced to mitigate feature collapse in Conformers by enhancing the nonlinearity in both the Multi-Head Attention and Feed-Forward Network components. The proposed architecture demonstrates superior efficiency and effectiveness, achieving state-of-the-art performance on the Lip Reading Sentences 2 (LRS2) and Lip Reading Sentences 3 (LRS3) datasets. An ablation study further validates the contributions of each component to the overall performance. The proposed model, which integrates Taylor-expanded attention and enhanced modules, exhibits robustness to noise and outperforms existing methods, offering a promising direction for the advancement of AVSR systems.