SimAM-UNet: A High-Fidelity Music Denoising Model in Complex Environments
摘要
With the development of mobile devices, the music recorded in complex environments often suffers from the interference of ambient noise that can damage the music quality and seriously degrade the listening experience. Extracting the noise feature in music accurately and filtering it effectively can be an effective way to alleviate these problems. However, compared with speech denoising, existing studies on music denoising are fewer and the applicable scenes are more homogeneous, which cannot meet the demands of complex environments in the real world. In our study, we proposed the SimAM-UNet model, which can filter out the environmental noise of music recorded in complex environments containing both single-source noise and multi-source noise. We developed the hybrid data augmentation mechanism in the noise augmenting and optimizing module to extend the pure noise audio dataset and improve the model generalization performance. We designed the noise representation learning module based on the SimAM mechanism to adaptively assign attention weights and improve the noise modeling capability. The experimental results show that the SimAM-UNet outperforms the state-of-the-art music denoising methods in both single-source noise and multi-source noise environments.