A Comparative Analysis of Advanced Deep Learning Techniques for Detecting Deepfake Audio
摘要
The rapid advancements in deep learning technologies have led to the generation of artificial audio, also known as audio deepfakes, becoming a significant concern for many sectors. These manipulations can be difficult to detect, and there is a need for more reliable detection models. The work investigates an approach for establishing a discrimination between “real” and “deepfake” audio by means of varying feature extraction methods including different machine learning and deep learning models. The research examines deep learning architectures such as convolutional neural networks (CNN), long short-term memory (LSTM), Res-Efficient CNN, XGBoost, Deep SONAR, etc. with a carefully selected dataset of real and synthetic audio samples. Research findings show the strengths and weaknesses of each model, and most of them have shown improvement in terms of detection accuracy with others. It demonstrated how effective each machine learning and deep learning model can be in detecting deepfake audio with some insights regarding the relative advantages of each approach.