<p>Existing deep acoustic echo cancellation (AEC) methods typically adopt single-domain modeling strategies in either the time domain or the frequency domain, making it difficult to simultaneously capture the dynamic correlations between the speech signal in both domains. Given speech signals’ significant time-frequency coupling characteristics, relying on just one domain for modelling fails to capture their joint structural features fully. To address this limitation, this paper introduces a novel end-to-end model, TFSPNet, which integrates a parallel time-frequency transformer module (PTFT) and a local perception attention module (LPAM). By combining these two modules, TFSPNet effectively captures both global contextual features and local structural information of the speech signal, greatly enhancing the capability for time-frequency joint modelling. This module combines a multi-head self-attention mechanism with a bidirectional GRU to enhance sequence modelling capabilities. At the same time, residual connections are introduced to improve the depth and precision of feature fusion. TFSPNet integrates the LPAM module within the PTFT to further enhance local structural modelling. This module employs channel-independent depthwise separable convolutions to extract spatial local attention maps, which are then fused with the original features element-wise, effectively emphasising key information regions and suppressing redundant interference. Experiments conducted on the AEC Challenge dataset show that, compared to the state-of-the-art model TF-DAD, TFSPNet achieves an 8.3% improvement in speech quality perception (PESQ), a short-term objective intelligibility (STOI) score of 0.95, and a 0.13 increase in the subjective mean opinion score (MOS), demonstrating its superior performance in challenging echo environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dual-branch time-frequency perception network for acoustic echo cancellation

  • Zhaodi Jiang,
  • Jian-Hong Wang,
  • Ji-Long He,
  • Xueting Li

摘要

Existing deep acoustic echo cancellation (AEC) methods typically adopt single-domain modeling strategies in either the time domain or the frequency domain, making it difficult to simultaneously capture the dynamic correlations between the speech signal in both domains. Given speech signals’ significant time-frequency coupling characteristics, relying on just one domain for modelling fails to capture their joint structural features fully. To address this limitation, this paper introduces a novel end-to-end model, TFSPNet, which integrates a parallel time-frequency transformer module (PTFT) and a local perception attention module (LPAM). By combining these two modules, TFSPNet effectively captures both global contextual features and local structural information of the speech signal, greatly enhancing the capability for time-frequency joint modelling. This module combines a multi-head self-attention mechanism with a bidirectional GRU to enhance sequence modelling capabilities. At the same time, residual connections are introduced to improve the depth and precision of feature fusion. TFSPNet integrates the LPAM module within the PTFT to further enhance local structural modelling. This module employs channel-independent depthwise separable convolutions to extract spatial local attention maps, which are then fused with the original features element-wise, effectively emphasising key information regions and suppressing redundant interference. Experiments conducted on the AEC Challenge dataset show that, compared to the state-of-the-art model TF-DAD, TFSPNet achieves an 8.3% improvement in speech quality perception (PESQ), a short-term objective intelligibility (STOI) score of 0.95, and a 0.13 increase in the subjective mean opinion score (MOS), demonstrating its superior performance in challenging echo environments.