<p>Synthetic speech detection systems have been proposed to bolster speaker verification systems against misinformation and speech spoofing. At present, the most effective methods are based on neural network models. The features extracted by different models often contain different information. Through the fusion of these models, their respective characteristics can be better combined. To fulfill this requirement, we propose the associative-discriminative fusion networks (ADFN) to fuse embeddings from different models. The ADFN can build the association between different models by the decomposition of a joint scatter matrix. And the decomposed matrices can be employed to compose the initialization transformation matrices of the input embeddings. With such structured initialization, the fusion model can start with a more suitable optimized location and achieve the better performance than each signal pre-fused model. The proposed ADFN method is evaluated on the Automatic Speaker Verification Spoofing and Countermeasures (ASVspoof) 2019 and 2021 datasets. The experimental results show that ADFN method achieves decent performance on the logical access (LA) sub-challenge of the two datasets, proving its practical applicability and effectiveness in fusing different types of embeddings.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Associative-Discriminative Fusion Networks for Synthetic Speech Detection

  • Kaijun Mai,
  • Chen Chen,
  • Yuhongxu Feng,
  • Deyun Chen

摘要

Synthetic speech detection systems have been proposed to bolster speaker verification systems against misinformation and speech spoofing. At present, the most effective methods are based on neural network models. The features extracted by different models often contain different information. Through the fusion of these models, their respective characteristics can be better combined. To fulfill this requirement, we propose the associative-discriminative fusion networks (ADFN) to fuse embeddings from different models. The ADFN can build the association between different models by the decomposition of a joint scatter matrix. And the decomposed matrices can be employed to compose the initialization transformation matrices of the input embeddings. With such structured initialization, the fusion model can start with a more suitable optimized location and achieve the better performance than each signal pre-fused model. The proposed ADFN method is evaluated on the Automatic Speaker Verification Spoofing and Countermeasures (ASVspoof) 2019 and 2021 datasets. The experimental results show that ADFN method achieves decent performance on the logical access (LA) sub-challenge of the two datasets, proving its practical applicability and effectiveness in fusing different types of embeddings.