Spoofing attacks, advanced text-to-speech (TTS) mechanisms, and voice conversion (VC) technologies pose significant security risks to Automatic Speaker Verification (ASV) systems by generating realistic-sounding synthetic speech and spoofed voices. To protect ASV systems from unauthorized access, it is essential to detect such synthetic speech and spoofed voices effectively. An Audio Spoof Detection (ASD) system typically comprises two main components: front-end feature extraction and back-end classification. In this paper, we introduce a novel approach to address this challenge by combining hand-crafted features with graph-based features during the front-end feature extraction stage. Specifically, we utilize Acoustic Ternary Pattern (ATP), Mel Frequency Cepstral Coefficient (MFCC), Gammatone Cepstral Coefficient (GTCC), and a new graph-based feature called Graph Frequency Cepstral Coefficient (GFCC). For the back-end classification, we employ Long Short-Term Memory (LSTM) networks. To evaluate the effectiveness of our proposed system, we develop six distinct configurations. The first three systems use ATP, MFCC, and GTCC features individually, while the remaining three systems combine GFCC with ATP, MFCC, and GTCC features, respectively. We conduct our experiments using the ASVspoof 2019 Physical Access (PA) evaluation datasets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Detection of Audio Spoofing Attack Using Integrated Hand-Crafted Features with LSTM

  • Nidhi Chakravarty,
  • Mohit Dua

摘要

Spoofing attacks, advanced text-to-speech (TTS) mechanisms, and voice conversion (VC) technologies pose significant security risks to Automatic Speaker Verification (ASV) systems by generating realistic-sounding synthetic speech and spoofed voices. To protect ASV systems from unauthorized access, it is essential to detect such synthetic speech and spoofed voices effectively. An Audio Spoof Detection (ASD) system typically comprises two main components: front-end feature extraction and back-end classification. In this paper, we introduce a novel approach to address this challenge by combining hand-crafted features with graph-based features during the front-end feature extraction stage. Specifically, we utilize Acoustic Ternary Pattern (ATP), Mel Frequency Cepstral Coefficient (MFCC), Gammatone Cepstral Coefficient (GTCC), and a new graph-based feature called Graph Frequency Cepstral Coefficient (GFCC). For the back-end classification, we employ Long Short-Term Memory (LSTM) networks. To evaluate the effectiveness of our proposed system, we develop six distinct configurations. The first three systems use ATP, MFCC, and GTCC features individually, while the remaining three systems combine GFCC with ATP, MFCC, and GTCC features, respectively. We conduct our experiments using the ASVspoof 2019 Physical Access (PA) evaluation datasets.