<p>Traditional methods for speech enhancement primarily focus on restoring amplitude features while neglecting phase information, which is equally critical for perceived quality. To address this issue, this paper proposes a dual-stream Conformer model that jointly processes complex and amplitude-domain features. The amplitude stream employs a mask to extract amplitude information, whereas the complex stream is responsible for capturing phase features. The model integrates Temporal Attention (TA), Dilated Convolution (DC), and Frequency Attention (FA) to extract both local and global speech features, with an attention-aware feature fusion (AFF) module for efficient dual-stream fusion, thereby enabling precise spectral estimation. The experimental results demonstrate that the proposed model significantly outperforms existing benchmarks on the VoiceBank + DEMAND dataset. Furthermore, ablation studies on the TIMIT dataset have been conducted to verify the contribution of each sub-module. The experimental results demonstrate that the proposed method achieves a 3.2% improvement over the latest methods on the dataset.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TSCME-Net: Two-Stream Conformer-Enhanced Network Leveraging Complex and Magnitude Spectrum Modeling for Noise-Robust Speech Enhancement

  • Zihao Feng,
  • Ying Gao,
  • Xinyu Guo,
  • Shifeng Ou

摘要

Traditional methods for speech enhancement primarily focus on restoring amplitude features while neglecting phase information, which is equally critical for perceived quality. To address this issue, this paper proposes a dual-stream Conformer model that jointly processes complex and amplitude-domain features. The amplitude stream employs a mask to extract amplitude information, whereas the complex stream is responsible for capturing phase features. The model integrates Temporal Attention (TA), Dilated Convolution (DC), and Frequency Attention (FA) to extract both local and global speech features, with an attention-aware feature fusion (AFF) module for efficient dual-stream fusion, thereby enabling precise spectral estimation. The experimental results demonstrate that the proposed model significantly outperforms existing benchmarks on the VoiceBank + DEMAND dataset. Furthermore, ablation studies on the TIMIT dataset have been conducted to verify the contribution of each sub-module. The experimental results demonstrate that the proposed method achieves a 3.2% improvement over the latest methods on the dataset.