Speech enhancement and separation in noisy reverberant environments are very challenging tasks. In this paper, we propose a speech enhancement and separation network, SESNet, for speech enhancement or speech separation in noisy reverberant environments, which is a multi-scale encoder-decoder architecture including a global-local feature extractor (GLFE). We also explored four kinds of Former blocks to be equipped in GLFE. We evaluate the performance of speech enhancement and speech separation on the VoiceBank+DEMAND and the WHAMR! datasets. The experimental results show that the SESNet has excellent performance for single- and multi-channel speech enhancement, and single-channel multi-speaker speech separation, keeping with a small model size.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SESNet: A Speech Enhancement and Separation Network in Noisy Reverberant Environments

  • Liusong Wang,
  • Yuan Gao,
  • Kaimin Cao,
  • Ying Hu

摘要

Speech enhancement and separation in noisy reverberant environments are very challenging tasks. In this paper, we propose a speech enhancement and separation network, SESNet, for speech enhancement or speech separation in noisy reverberant environments, which is a multi-scale encoder-decoder architecture including a global-local feature extractor (GLFE). We also explored four kinds of Former blocks to be equipped in GLFE. We evaluate the performance of speech enhancement and speech separation on the VoiceBank+DEMAND and the WHAMR! datasets. The experimental results show that the SESNet has excellent performance for single- and multi-channel speech enhancement, and single-channel multi-speaker speech separation, keeping with a small model size.