Current synthetic speech detection methods often overlook the emotional distinctions between synthetic and real speech. To comprehensively leverage emotional information and deep features to further enhance the accuracy of distinguishing synthetic speeches, this paper proposes an emotion-aware synthetic speech detection method by fusing multi-level different features. Initially, a multi-scale emotional feature extraction model is devised to extract emotional features from 3-D log-Mels. Considering that the emotional features learned by a single model may not be comprehensive enough, the emotion2vec model is utilized to automatically learn additional emotional features from the temporal domain of the original speech. Subsequently, by employing a co-attention network, the extracted two kinds of emotional features are both incorporated into the deep features learned by the end-to-end model Rawnet2, thereby obtaining highly sensitive and weighted deep features. Finally, the extracted emotional features and the weighted deep features are fused to classify synthetic speech from real speech. Experimental results on several datasets demonstrate that the proposed method outperforms baseline methods, with a detection accuracy exceeding 0.97.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Emotion-Aware Synthetic Speech Detection with Multi-level Fusion Features

  • Lingyun Xiang,
  • Zhiguo Zhou,
  • Yangfan Liu

摘要

Current synthetic speech detection methods often overlook the emotional distinctions between synthetic and real speech. To comprehensively leverage emotional information and deep features to further enhance the accuracy of distinguishing synthetic speeches, this paper proposes an emotion-aware synthetic speech detection method by fusing multi-level different features. Initially, a multi-scale emotional feature extraction model is devised to extract emotional features from 3-D log-Mels. Considering that the emotional features learned by a single model may not be comprehensive enough, the emotion2vec model is utilized to automatically learn additional emotional features from the temporal domain of the original speech. Subsequently, by employing a co-attention network, the extracted two kinds of emotional features are both incorporated into the deep features learned by the end-to-end model Rawnet2, thereby obtaining highly sensitive and weighted deep features. Finally, the extracted emotional features and the weighted deep features are fused to classify synthetic speech from real speech. Experimental results on several datasets demonstrate that the proposed method outperforms baseline methods, with a detection accuracy exceeding 0.97.