Speaker Extraction with Verification of Present and Absent Target Speakers
摘要
Target speaker extraction (TSE) models are expected to extract the target speech from a cocktail party mixture signal. When only trained with present target speaker samples (PT), these models output noise in the absence of the target speaker (AT). One may enhance the TSE quality by providing the information about the PT and AT. However, the detection of the target speaker is not perfect. In this paper, we propose a new model, TSEV, which performs target speaker extraction and speaker verification simultaneously. The TSEV model outputs an extracted speech and generates two speaker embeddings per inference to detect the target speaker. By sharing the speaker encoder and low-level modules, the speaker verification task can be performed in low signal-to-noise ratio scenarios. We train the TSEV model on multi-talker PT and AT conditions with fully overlapped speech. Experiments verify the superiority of jointly performing two tasks in the proposed model. The TSEV model achieves better verification performance without degrading the extraction performance compared with the baseline.