The convolution-augmented Transformer (Conformer) has emerged as a promising architecture in the field of speech enhancement (SE), capable of capturing both local and global dependencies in speech signals. Recent studies have shown that replacing the self-attention mechanism with unparameterized token mixers yields results that are similar to or superior to those of traditional Transformers, with significantly reduced computational and memory requirements. This has led to the development of the Metaformer architecture, which consists of normalization, residual connections, channel mixers, and token mixers. The fundamental architecture comprising these components is primarily responsible for the strong performance of the models, rather than the specific choice of token mixer. So far, Metaformers have mainly been tested in the domains of language modeling, document-grounded dialogue generation, image classification, object detection and instance segmentation. In this paper, we transfer the Metaformer concept to the field of speech enhancement. We combine the Metaformer and Conformer architectures, resulting in the MetaConformer model, which utilizes token mixers instead of self-attention. By substituting the Conformer block in a state-of-the-art speech enhancement model with our MetaConformer concept, we demonstrate that even the simplest token mixers can achieve 95% of the PESQ score of the vanilla model on the VoiceBank+DEMAND dataset, while requiring only 57% of the training time and 63% less memory. The introduced MetaConformer could serve as a baseline for future speech enhancement models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exploring MetaConformer for Speech Enhancement

  • Lukas Förner,
  • Maximilian Dauner

摘要

The convolution-augmented Transformer (Conformer) has emerged as a promising architecture in the field of speech enhancement (SE), capable of capturing both local and global dependencies in speech signals. Recent studies have shown that replacing the self-attention mechanism with unparameterized token mixers yields results that are similar to or superior to those of traditional Transformers, with significantly reduced computational and memory requirements. This has led to the development of the Metaformer architecture, which consists of normalization, residual connections, channel mixers, and token mixers. The fundamental architecture comprising these components is primarily responsible for the strong performance of the models, rather than the specific choice of token mixer. So far, Metaformers have mainly been tested in the domains of language modeling, document-grounded dialogue generation, image classification, object detection and instance segmentation. In this paper, we transfer the Metaformer concept to the field of speech enhancement. We combine the Metaformer and Conformer architectures, resulting in the MetaConformer model, which utilizes token mixers instead of self-attention. By substituting the Conformer block in a state-of-the-art speech enhancement model with our MetaConformer concept, we demonstrate that even the simplest token mixers can achieve 95% of the PESQ score of the vanilla model on the VoiceBank+DEMAND dataset, while requiring only 57% of the training time and 63% less memory. The introduced MetaConformer could serve as a baseline for future speech enhancement models.