Current speech spoofing detection technologies often perform poorly against unknown attacks due to insufficient generalization and inadequate feature extraction, particularly in cross-dataset scenarios. In this paper, an end-to-end spoofing detection model named RawSSLNet is proposed to significantly enhance detection performance and generalization capabilities. The model improves on the end-to-end RawNet2 by incorporating a self-supervised pre-trained model, wav2vec 2.0, to extract more robust acoustic features, integrating a squeeze-and-excitation operation, and employing feature map scaling (FMS) in 2D space to further enhance these features. Additionally, a mixture of additive noise is introduced to enhance the data. Compared to the baseline model, the proposed model achieves significantly lower equal error rate (EER) and minimum tandem detection cost function (min t-DCF) on the ASVspoof2019, ASVspoof2021, and In-the-Wild (ITW) datasets. The results clearly indicate that the improved model can extract acoustic features with greater accuracy and comprehensiveness while consistently maintaining superior performance in challenging cross-dataset detection scenarios.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Spoofing Speech Detection Method Based on Self-supervised Front End and Feature Enhancement

  • Boyan Guo,
  • Guangcun Wei,
  • Chunyu Meng,
  • Jianan Liu,
  • Chengde Zhang

摘要

Current speech spoofing detection technologies often perform poorly against unknown attacks due to insufficient generalization and inadequate feature extraction, particularly in cross-dataset scenarios. In this paper, an end-to-end spoofing detection model named RawSSLNet is proposed to significantly enhance detection performance and generalization capabilities. The model improves on the end-to-end RawNet2 by incorporating a self-supervised pre-trained model, wav2vec 2.0, to extract more robust acoustic features, integrating a squeeze-and-excitation operation, and employing feature map scaling (FMS) in 2D space to further enhance these features. Additionally, a mixture of additive noise is introduced to enhance the data. Compared to the baseline model, the proposed model achieves significantly lower equal error rate (EER) and minimum tandem detection cost function (min t-DCF) on the ASVspoof2019, ASVspoof2021, and In-the-Wild (ITW) datasets. The results clearly indicate that the improved model can extract acoustic features with greater accuracy and comprehensiveness while consistently maintaining superior performance in challenging cross-dataset detection scenarios.