<p>Single channel speech enhancement has been extensively researched by employing real-valued neural networks in time–frequency (T/F) domain. Although the input and target in the T/F domain are naturally complex, a full complex model can effectively learn a complex sequence and learn feature representations. Furthermore, the phase component, which is crucial for the perceptual quality, has been demonstrated to be learnable alongside the magnitude component, even in the presence of noise, through methods such as complex masking (CM) or complex spectral mapping (CSM). Several recent studies concentrate on either CM or CSM, without considering their performance limits. To tackle the aforementioned challenges, we propose a comprehensive solution known as the D2-Transformer (dual-path dual-decoder multi-axial transformer network). This approach utilizes CM and CSM to enhance single channel speech. In this paper, we introduce the lightweight multi-axial transformer (MA Transformer) that extracts feature efficiently along both temporal and frequency axes, while requiring minimal computational overhead. The Transformer’s ability to utilize local features effectively is improved by its time/frequency multi-DCONV head self-attention block (T/F-M-DCHSA). Furthermore, we implement a block called the frequency prompt block, which guides varying levels of degradation in speech signals into recovering frequency features. This allows for effective modeling of sequences in the complex T/F domain. To enhance the representation of time–frequency features in both the encoder as well as decoders, we employ a dual-path learning approach. This involves utilizing complex dilated-convolutions to capture time dependencies and complex feedforward sequential memory networks to address frequency recurrence. Furthermore, we enhance the performance limits of CM and CSM by integrating the advantages of both training targets within a unified learning context. As a result, D2-Transformer maximizes the benefits of complex operations, dual-path framework, and joint-training targets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Single Channel Speech Enhancement using a Complex Dual-Path Multi Axial Transformer with Frequency Prompt

  • Chaitanya Jannu,
  • Manaswini Burra,
  • Sunny Dayal Vanambathina,
  • Veeraswamy Parisae,
  • Chinta Venkata Murali Krishna,
  • G. L. Madhumati

摘要

Single channel speech enhancement has been extensively researched by employing real-valued neural networks in time–frequency (T/F) domain. Although the input and target in the T/F domain are naturally complex, a full complex model can effectively learn a complex sequence and learn feature representations. Furthermore, the phase component, which is crucial for the perceptual quality, has been demonstrated to be learnable alongside the magnitude component, even in the presence of noise, through methods such as complex masking (CM) or complex spectral mapping (CSM). Several recent studies concentrate on either CM or CSM, without considering their performance limits. To tackle the aforementioned challenges, we propose a comprehensive solution known as the D2-Transformer (dual-path dual-decoder multi-axial transformer network). This approach utilizes CM and CSM to enhance single channel speech. In this paper, we introduce the lightweight multi-axial transformer (MA Transformer) that extracts feature efficiently along both temporal and frequency axes, while requiring minimal computational overhead. The Transformer’s ability to utilize local features effectively is improved by its time/frequency multi-DCONV head self-attention block (T/F-M-DCHSA). Furthermore, we implement a block called the frequency prompt block, which guides varying levels of degradation in speech signals into recovering frequency features. This allows for effective modeling of sequences in the complex T/F domain. To enhance the representation of time–frequency features in both the encoder as well as decoders, we employ a dual-path learning approach. This involves utilizing complex dilated-convolutions to capture time dependencies and complex feedforward sequential memory networks to address frequency recurrence. Furthermore, we enhance the performance limits of CM and CSM by integrating the advantages of both training targets within a unified learning context. As a result, D2-Transformer maximizes the benefits of complex operations, dual-path framework, and joint-training targets.