<p>Traditional multi-stage speech enhancement methods based on time–frequency (T–F) domain target decoupling could improve the amplitude and phase compensation issues. However, due to the lack of focus on the relationship between multi-scale information and inter-structure within the T–F features, their enhancement performance remain limited in real applications. To address this, a two-stage speech enhancement method based on dual-branch multi-scale T–F attention (DMTFA) is proposed. First, a parallel DMTFA module with fewer parameters is designed as the feature processing module of the encoders and decoders, which primarily focuses on the positional information of each layer’s features while extracting multi-scale features in parallel. Second, an exponential weighting operation is applied to optimize the temporal convolutional modules, strengthening the correlation between T–F bins. Additionally, the pooling strategies along the time dimension is improved using a sliding window, converting the non-causal model into a causal model, thus extending the practical applicability of the proposed method. Finally, multiple sets of experiments are conducted to validate our causal and non-causal models. The results show that these models achieve higher speech enhancement performance with fewer parameters and lower complexity.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Two-Stage Speech Enhancement Based on Dual-Branch Multi-scale Time–Frequency Attention

  • Zekang Qi,
  • Yangjie Wei,
  • Bingbing Wang,
  • Zhuangzhuang Wang

摘要

Traditional multi-stage speech enhancement methods based on time–frequency (T–F) domain target decoupling could improve the amplitude and phase compensation issues. However, due to the lack of focus on the relationship between multi-scale information and inter-structure within the T–F features, their enhancement performance remain limited in real applications. To address this, a two-stage speech enhancement method based on dual-branch multi-scale T–F attention (DMTFA) is proposed. First, a parallel DMTFA module with fewer parameters is designed as the feature processing module of the encoders and decoders, which primarily focuses on the positional information of each layer’s features while extracting multi-scale features in parallel. Second, an exponential weighting operation is applied to optimize the temporal convolutional modules, strengthening the correlation between T–F bins. Additionally, the pooling strategies along the time dimension is improved using a sliding window, converting the non-causal model into a causal model, thus extending the practical applicability of the proposed method. Finally, multiple sets of experiments are conducted to validate our causal and non-causal models. The results show that these models achieve higher speech enhancement performance with fewer parameters and lower complexity.