Enhancing ASV Security: An Integrated Ensemble Detector for Logical Access and DeepFake Attacks
摘要
Automated Speaker Verification (ASV) systems play a crucial role in securing voice-based authentication for voice-driven devices (VDDs) like Google Home and Amazon Alexa extensively used in consumer IoT applications. They are susceptible to various voice-based attacks, including Logical Access (LA), Physical Access (PA), and DeepFake (DF) attacks, posing significant security challenges. Conventional countermeasures have been outpaced by rapid advances in creating sophisticated audio spoofs. We propose a voice spoofing detection method using two distinct audio feature extraction approaches, capturing dynamic speech attributes and algorithmic artifacts. An integrated ensemble approach, combines features from raw audio waveforms and cepstral representations from spectrograms of the audio signal to make the final decision. Visual representations of the spectrograms are learned using a fine-tuned series of Squeeze-and-Excitation (SE) convolution blocks. The concatenated features are further passed through a pair of dense layers activated with softmax. Experimentation on the LA and DF tracks of ASVspoof 2021 dataset showcases the effectiveness of our proposed spoofing detection system, achieving an accuracy of 95.84% and an Equal Error Rate of 12.43% on the LA task, and an accuracy of 97.00% and an Equal Error Rate of 21.68% on the DF task. Our method demonstrates potential in enhancing the security of speaker verification systems.