<p>This paper tries to address “the cocktail party problem”, focusing on multi-talker speech separation and recognition-tasks where current deep learning methods still lag behind human performance. Inspired by the human auditory system, which discerns individual speakers in noisy environments by understanding context, we propose a novel architecture that enhances monaural speech separation and multi-talker recognition by leveraging contextual embeddings directly from mixed speech. We introduce a single-talker teaching multi-talker knowledge distillation framework for learning these embeddings, enabling accurate contextual predictions for each speaker. Additionally, we explore integrating these embeddings into separation and recognition systems, supported by advanced optimization techniques like embedding sampling and two-stage training. Evaluations on the WSJ0-2mix benchmark demonstrate significant performance improvements. Analysis shows that the learned embeddings effectively capture and segregate contextual information, making them highly beneficial for separating and recognizing speech in multi-talker scenarios.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Contextual understanding with contextual embeddings for multi-talker speech separation and recognition in a cocktail party

  • Yanmin Qian,
  • Chenda Li,
  • Wangyou Zhang,
  • Shaoxiong Lin

摘要

This paper tries to address “the cocktail party problem”, focusing on multi-talker speech separation and recognition-tasks where current deep learning methods still lag behind human performance. Inspired by the human auditory system, which discerns individual speakers in noisy environments by understanding context, we propose a novel architecture that enhances monaural speech separation and multi-talker recognition by leveraging contextual embeddings directly from mixed speech. We introduce a single-talker teaching multi-talker knowledge distillation framework for learning these embeddings, enabling accurate contextual predictions for each speaker. Additionally, we explore integrating these embeddings into separation and recognition systems, supported by advanced optimization techniques like embedding sampling and two-stage training. Evaluations on the WSJ0-2mix benchmark demonstrate significant performance improvements. Analysis shows that the learned embeddings effectively capture and segregate contextual information, making them highly beneficial for separating and recognizing speech in multi-talker scenarios.