Conformer-Based Audio Visual Speech Recognition with Taylor Attention
摘要
Audio Visual Speech Recognition (AVSR) has witnessed significant advancements by deploying sophisticated neural networks, particularly Convolutional Neural Networks (CNNs) and Transformers, to improve automatic speech recognition, lip-reading, and audio-visual scene understanding. However, challenges persist in effectively integrating Transformers—known for their ability to capture global context—into AVSR systems due to their substantial computational demands. This paper introduces a novel approach utilizing a modified version of the Conformer architecture, termed the Taylor-Expanded Linearized Conformer, designed to address these challenges. The proposed method retains the capability to model long-range dependencies while ensuring computational complexity remains linear. Additionally, a new architecture is introduced to mitigate feature collapse in Conformers by enhancing the nonlinearity in both the Multi-Head Attention and Feed-Forward Network components. The proposed architecture demonstrates superior efficiency and effectiveness, achieving state-of-the-art performance on the Lip Reading Sentences 2 (LRS2) and Lip Reading Sentences 3 (LRS3) datasets. An ablation study further validates the contributions of each component to the overall performance. The proposed model, which integrates Taylor-expanded attention and enhanced modules, exhibits robustness to noise and outperforms existing methods, offering a promising direction for the advancement of AVSR systems.