Exploring Using Contrastive Learning for Improving BSRNN-Based Speech Enhancement
摘要
BSRNN, a band-split RNN model designed for tasks such as speech enhancement and source separation, has demonstrated outstanding performance in recent research. Meanwhile, contrastive learning, as a self-supervised learning framework, has proven effective in guiding networks to learn speech representations with broad potential across domains. This paper explores integrating contrastive learning techniques into BSRNN to optimize its speech enhancement capabilities. Specifically, we design various contrastive losses that complement the supervised loss to jointly train the model. To improve the shallow acoustic representations generated by BSRNN’s band-split module, we first investigate two types of noisy-representation contrastive learning: one applied between full-band discrete and band-split continuous representations, and another between band-split discrete and continuous representations. Additionally, we propose a noisy-clean representation contrastive learning method to enhance the model’s denoising ability. By enabling mutual reinforcement between contrastive and supervised losses during training, the proposed approach significantly improves performance without introducing extra parameters or computational cost during inference. Experiments on the VoiceBank+DEMAND dataset demonstrate that all the proposed methods outperform the original BSRNN across various evaluation metrics, highlighting more efficient and effective solutions in BSRNN-based speech enhancement tasks.