<p>Speech enhancement improves speech quality by mitigating noise, dereverberation, and echo. Existing methods face challenges in amplitude-phase compensation and temporal-frequency feature modeling, while also suffering from high computational complexity. To address these issues, a gated dense encoder-decoder architecture with a two-stage Conformer, abbreviated as GD-Conformer, is proposed for monaural speech enhancement. It integrates a gated dense encoder, a two-stage residual Conformer module, a mask decoder, and a complex decoder. The gated dense encoder consists of two parts: a dilated dense convolution and a gated convolution, where the former captures both global and local dependency features, while the latter refines these distinct features accordingly. The two-stage residual Conformer focuses on the time-frequency dependence of speech, it also reduces the computational complexity. The mask decoder estimates the mask for the input magnitude, enhancing the effectiveness of magnitude features. And the complex decoder refines the real and imaginary components of the input, compensating for phase features to accurately reconstruct the clean speech signal. The outcomes of experiments conducted on the public dataset VoiceBank+DEMAND and DNS Challenge 2020 demonstrate that, compared with those state-of-the-art methods, the proposed GD-Conformer achieves comparable performance in terms of denoising and generalization with fewer parameters and lower computation complexity.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

GD-Conformer: a Conformer-based gated dense encoder-decoder for monaural speech enhancement

  • Gengzangcuomao,
  • Heming Huang,
  • Feipeng Da

摘要

Speech enhancement improves speech quality by mitigating noise, dereverberation, and echo. Existing methods face challenges in amplitude-phase compensation and temporal-frequency feature modeling, while also suffering from high computational complexity. To address these issues, a gated dense encoder-decoder architecture with a two-stage Conformer, abbreviated as GD-Conformer, is proposed for monaural speech enhancement. It integrates a gated dense encoder, a two-stage residual Conformer module, a mask decoder, and a complex decoder. The gated dense encoder consists of two parts: a dilated dense convolution and a gated convolution, where the former captures both global and local dependency features, while the latter refines these distinct features accordingly. The two-stage residual Conformer focuses on the time-frequency dependence of speech, it also reduces the computational complexity. The mask decoder estimates the mask for the input magnitude, enhancing the effectiveness of magnitude features. And the complex decoder refines the real and imaginary components of the input, compensating for phase features to accurately reconstruct the clean speech signal. The outcomes of experiments conducted on the public dataset VoiceBank+DEMAND and DNS Challenge 2020 demonstrate that, compared with those state-of-the-art methods, the proposed GD-Conformer achieves comparable performance in terms of denoising and generalization with fewer parameters and lower computation complexity.