The Attention-Based Fusion of Master-Auxiliary Network for Speech Enhancement
摘要
To address the issue of insufficient learning capability in a single network, a master-auxiliary network incorporating fusion attention mechanisms is proposed for speech enhancement. In the master network, different-shaped convolution kernels are initially employed to extract time-domain and frequency-domain information separately, resulting in an enhanced CNN. Subsequently, this enhanced CNN is paralleled with a BiGRU integrated with a fusion multi-head attention mechanism to mitigate the CNN’s inability to utilize global information. Moreover, attention mechanisms are employed to focus on the positional and channel dimension information within the CNN. Additionally, to address the issues of oversmoothed outputs and insufficient perception of high-frequency information in existing networks, this paper proposes an Exponential Target Exaggerated Mean Squared Error (ETMSE) loss function. In the auxiliary network, by training on the residual between the output of the master network and the learning target, the obtained residuals are compensated back to the master network, thereby reducing the loss of some speech information during the learning process. Finally, experiments are conducted on multiple speeches with different background noises and various signal-to-noise ratios. Compared to the traditional CNN algorithm, the enhanced speech PESQ mean value increased by approximately 0.42, validating the effectiveness of the proposed algorithm.