Robustness against channel variability remains a formidable challenge in audio anti-spoofing for speaker verification systems. Channel effects can significantly degrade the performance of countermeasure systems, making them susceptible to spoofing attacks. To address this challenge, we present a comprehensive approach that integrates channel-robust preprocessing with advanced graph-based neural networks to enhance detection reliability. Raw audio waveforms are preprocessed with data augmentation to simulate diverse acoustic conditions, and encoded using a modified RawNet2-based encoder to extract critical features. An adaptive graph module processes these features into spectral and temporal graphs. Our proposed method dynamically combines these graphs using a Heterogeneous Stacking Graph Attention Layer (HS-GAL), facilitating deeper integration and processing of audio data. The Max Graph Operation (MGO) further refines feature selection, crucial for identifying spoofed content. Additionally, our model incorporates adversarial and multi-task learning strategies, significantly enhancing its generalization capabilities across various datasets. Experimental results demonstrate that our approach reduces the Equal Error Rate (EER) by over 20% and the minimum tandem detection cost function (min t-DCF) by 25% relative to the current state-of-the-art, substantiating its efficacy in improving the security of speaker verification systems against channel-induced vulnerabilities.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Channel Robust Strategies with Data Augmentation for Audio Anti-spoofing

  • Sardor Mamarasulov,
  • Yang Li,
  • Changbo Wang

摘要

Robustness against channel variability remains a formidable challenge in audio anti-spoofing for speaker verification systems. Channel effects can significantly degrade the performance of countermeasure systems, making them susceptible to spoofing attacks. To address this challenge, we present a comprehensive approach that integrates channel-robust preprocessing with advanced graph-based neural networks to enhance detection reliability. Raw audio waveforms are preprocessed with data augmentation to simulate diverse acoustic conditions, and encoded using a modified RawNet2-based encoder to extract critical features. An adaptive graph module processes these features into spectral and temporal graphs. Our proposed method dynamically combines these graphs using a Heterogeneous Stacking Graph Attention Layer (HS-GAL), facilitating deeper integration and processing of audio data. The Max Graph Operation (MGO) further refines feature selection, crucial for identifying spoofed content. Additionally, our model incorporates adversarial and multi-task learning strategies, significantly enhancing its generalization capabilities across various datasets. Experimental results demonstrate that our approach reduces the Equal Error Rate (EER) by over 20% and the minimum tandem detection cost function (min t-DCF) by 25% relative to the current state-of-the-art, substantiating its efficacy in improving the security of speaker verification systems against channel-induced vulnerabilities.