Detection of Speech Spoofing Based on Dense Convolutional Network
摘要
In recent years, the rapid development of voice synthesis technologies has led to an increasing concern about the abuse of fake human voices for malicious purposes, such as deepfake audio, spam calls and social engineering attacks. This paper proposes a novel deep learning-based model to effectively identify counterfeit human voices generated by various voice synthesis algorithms. The proposed model employs a combination of Dense-Style Network to capture both spectral and temporal features of human speech. The model is extensively evaluated on ASVspoof 2019 datasets. The experimental results indicate that our model achieves competitive performance compared to existing methods and has a certain degree of anti-compression ability. In addition, anti-compression research was conducted to investigate the recognition performance of the model in response to compressed speech. Our findings pave the way for further research in combating against the misuse of artificially generated human voices and sound authenticity verification in general.