Speech Deepfake Recognition with Data-Efficient Vision Transformers
摘要
Multimedia deepfake detection is an increasingly important task due to the rapid development of generative models and broad possibilities of malicious usage. In this paper, we propose a method of speech deepfake detection using data-efficient image transformers (DeiT) and spectrogram representations. We compare our solution with existing approaches and show that our solution improves on the previously reported state of the art results. Our results show that data-efficient image transformers work well for several methods of generating synthetic speech. Furthermore, our proposed approach is applicable in low-compute scenarios, as it requires only a single consumer grade GPU.