<p>Target speaker extraction (TSE) models are expected to extract the target speech from a cocktail party mixture signal. When only trained with present target speaker samples (PT), these models output noise in the absence of the target speaker (AT). One may enhance the TSE quality by providing the information about the PT and AT. However, the detection of the target speaker is not perfect. In this paper, we propose a new model, TSEV, which performs target speaker extraction and speaker verification simultaneously. The TSEV model outputs an extracted speech and generates two speaker embeddings per inference to detect the target speaker. By sharing the speaker encoder and low-level modules, the speaker verification task can be performed in low signal-to-noise ratio scenarios. We train the TSEV model on multi-talker PT and AT conditions with fully overlapped speech. Experiments verify the superiority of jointly performing two tasks in the proposed model. The TSEV model achieves better verification performance without degrading the extraction performance compared with the baseline.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Speaker Extraction with Verification of Present and Absent Target Speakers

  • Ke Zhang,
  • Marvin Borsdorf,
  • Tianchi Liu,
  • Shuai Wang,
  • Yangjie Wei,
  • Haizhou Li

摘要

Target speaker extraction (TSE) models are expected to extract the target speech from a cocktail party mixture signal. When only trained with present target speaker samples (PT), these models output noise in the absence of the target speaker (AT). One may enhance the TSE quality by providing the information about the PT and AT. However, the detection of the target speaker is not perfect. In this paper, we propose a new model, TSEV, which performs target speaker extraction and speaker verification simultaneously. The TSEV model outputs an extracted speech and generates two speaker embeddings per inference to detect the target speaker. By sharing the speaker encoder and low-level modules, the speaker verification task can be performed in low signal-to-noise ratio scenarios. We train the TSEV model on multi-talker PT and AT conditions with fully overlapped speech. Experiments verify the superiority of jointly performing two tasks in the proposed model. The TSEV model achieves better verification performance without degrading the extraction performance compared with the baseline.