Deep learning innovations have been pivotal in advancing the field of source separation, especially in challenging environments such as distinguishing parallel conversations at a party and utilizing virtual assistants or making a call in crowded circumstance. This study introduces a novel contrastive learning-based methodology for separating two speech sources from background music in party-like situations. Given inputs that include both music and speech sources, our method applies an encoder-separator-decoder framework to be a background system for separating audio by leveraging contrastive learning. Our research adopts a multi-negative approach, employing tuplet loss and N-pair loss functions, to focus on individual audio sources. Additionally, our method utilized a specially curated dataset for party scenarios, composed of Royalty free audio from Kaggle, the Librimix dataset, and the Musan dataset. The outcomes demonstrate remarkable results, achieving a Scale-Invariant Signal-to-Noise Ratio improvements (SI-SNRi) of 14.81 dB for separating speaker 1, speaker 2 and music from the mixture, a SI-SNRi of 18.33 dB for separating both speakers from the mixture, and a SI-SNRi of 7.78 dB for separating music alone from the mixture, evidencing the effectiveness of our approach in disentangling individual audio components from complex acoustic mixtures.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Separating Party Conversation by Applying Contrastive Learning Methodology

  • Anandakumar Singaravelan,
  • Jia-Lien Hsu

摘要

Deep learning innovations have been pivotal in advancing the field of source separation, especially in challenging environments such as distinguishing parallel conversations at a party and utilizing virtual assistants or making a call in crowded circumstance. This study introduces a novel contrastive learning-based methodology for separating two speech sources from background music in party-like situations. Given inputs that include both music and speech sources, our method applies an encoder-separator-decoder framework to be a background system for separating audio by leveraging contrastive learning. Our research adopts a multi-negative approach, employing tuplet loss and N-pair loss functions, to focus on individual audio sources. Additionally, our method utilized a specially curated dataset for party scenarios, composed of Royalty free audio from Kaggle, the Librimix dataset, and the Musan dataset. The outcomes demonstrate remarkable results, achieving a Scale-Invariant Signal-to-Noise Ratio improvements (SI-SNRi) of 14.81 dB for separating speaker 1, speaker 2 and music from the mixture, a SI-SNRi of 18.33 dB for separating both speakers from the mixture, and a SI-SNRi of 7.78 dB for separating music alone from the mixture, evidencing the effectiveness of our approach in disentangling individual audio components from complex acoustic mixtures.