Deepfake audio is synthetic voice generated using sophisticated artificial intelligence. It mimics human speech patterns, intonation, and voice characteristics. The deepfake audio undermines trust, authenticity of information, and security by spreading misinformation and engaging in social engineering attacks. To mitigate the risk, we proposed an optimized and generalized deep learning architecture to classify authentic human audio and AI-generated cloning voices. We used LSTM and its variant, BiLSTM, with optimal optimization and regularization on the benchmark dataset. The authentic recorded audio is used as authentic and then transformed using AI. We approach it as a binary classification problem and use different statistical analyses to distinguish variation in the distribution of temporal characteristics. The proposed architecture successfully captures audio data’s local spectrum information and long-range temporal relationships. We used different performance evaluation matrices to check the performance, and overall, we got around 99% accuracy. This framework can handle a variety of audio sources and is resistant to advanced deep-fake techniques.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

UniAudioGuard: Universal DL Architecture for AI-Generated Deepfake Audio Classification

  • Om Prakash

摘要

Deepfake audio is synthetic voice generated using sophisticated artificial intelligence. It mimics human speech patterns, intonation, and voice characteristics. The deepfake audio undermines trust, authenticity of information, and security by spreading misinformation and engaging in social engineering attacks. To mitigate the risk, we proposed an optimized and generalized deep learning architecture to classify authentic human audio and AI-generated cloning voices. We used LSTM and its variant, BiLSTM, with optimal optimization and regularization on the benchmark dataset. The authentic recorded audio is used as authentic and then transformed using AI. We approach it as a binary classification problem and use different statistical analyses to distinguish variation in the distribution of temporal characteristics. The proposed architecture successfully captures audio data’s local spectrum information and long-range temporal relationships. We used different performance evaluation matrices to check the performance, and overall, we got around 99% accuracy. This framework can handle a variety of audio sources and is resistant to advanced deep-fake techniques.