Associative-Discriminative Fusion Networks for Synthetic Speech Detection
摘要
Synthetic speech detection systems have been proposed to bolster speaker verification systems against misinformation and speech spoofing. At present, the most effective methods are based on neural network models. The features extracted by different models often contain different information. Through the fusion of these models, their respective characteristics can be better combined. To fulfill this requirement, we propose the associative-discriminative fusion networks (ADFN) to fuse embeddings from different models. The ADFN can build the association between different models by the decomposition of a joint scatter matrix. And the decomposed matrices can be employed to compose the initialization transformation matrices of the input embeddings. With such structured initialization, the fusion model can start with a more suitable optimized location and achieve the better performance than each signal pre-fused model. The proposed ADFN method is evaluated on the Automatic Speaker Verification Spoofing and Countermeasures (ASVspoof) 2019 and 2021 datasets. The experimental results show that ADFN method achieves decent performance on the logical access (LA) sub-challenge of the two datasets, proving its practical applicability and effectiveness in fusing different types of embeddings.