Separating Party Conversation by Applying Contrastive Learning Methodology
摘要
Deep learning innovations have been pivotal in advancing the field of source separation, especially in challenging environments such as distinguishing parallel conversations at a party and utilizing virtual assistants or making a call in crowded circumstance. This study introduces a novel contrastive learning-based methodology for separating two speech sources from background music in party-like situations. Given inputs that include both music and speech sources, our method applies an encoder-separator-decoder framework to be a background system for separating audio by leveraging contrastive learning. Our research adopts a multi-negative approach, employing tuplet loss and N-pair loss functions, to focus on individual audio sources. Additionally, our method utilized a specially curated dataset for party scenarios, composed of Royalty free audio from Kaggle, the Librimix dataset, and the Musan dataset. The outcomes demonstrate remarkable results, achieving a Scale-Invariant Signal-to-Noise Ratio improvements (SI-SNRi) of 14.81 dB for separating speaker 1, speaker 2 and music from the mixture, a SI-SNRi of 18.33 dB for separating both speakers from the mixture, and a SI-SNRi of 7.78 dB for separating music alone from the mixture, evidencing the effectiveness of our approach in disentangling individual audio components from complex acoustic mixtures.