In recent years, the advancements made in Artificial intelligence (AI) image generation techniques have seen a rise in realistic deepfakes used for nefarious purposes. Such images or videos have become so realistic that it is becoming increasingly difficult to tell them apart from real images, causing concerns on societal impacts. In this work, deep learning models with pre-trained backbone model that consider the nature of face deepfakes in the wild, were developed. Also, newer variants of the pre-trained EfficientNET architecture such as EfficientNET, V2S, and V2M were used as the backbone of the face deepfake detection models in addition to the older variants seen in literature such as EfficientNET B3 and B5 to compare their performance against face deepfakes. More importantly, images generated by the latest StyleGAN generation technique called StyleGAN3 were included and applied in this work. The inclusion of StyleGAN3 generated images in a deep-fake detection model has not been observed in literature yet, therefore this paper can act as a starting point for reference on how pre-trained models such as EfficientNET perform against StyleGAN3. The results show the best accuracy rate is 92.77% achieved by the EfficientNET B3 backbone model, beating newer EfficientNET variants for the non-StyleGAN3 test set. However, when tested against a small pre-processed StyleGAN3 test set prepared separately, the best accuracy score achieved was only 64.29%. In addition, this was achieved by EfficientNET V2M which scored only 88.60% accuracy. The results demonstrate that even pretrained backbone-based detectors with very high performance against images generated by a variety of generation techniques including StyleGAN3 and some “in the wild” scenarios considered, struggle to retain their accuracy on a small set of pre-processed StyleGAN3 generated images.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

EfficientNET-Based Deepfake Detection: Performance Analysis Against StyleGAN3 and In-the-Wild Scenarios

  • Fatimah Meraj,
  • Essa Q. Shahra,
  • Shadi Basurra

摘要

In recent years, the advancements made in Artificial intelligence (AI) image generation techniques have seen a rise in realistic deepfakes used for nefarious purposes. Such images or videos have become so realistic that it is becoming increasingly difficult to tell them apart from real images, causing concerns on societal impacts. In this work, deep learning models with pre-trained backbone model that consider the nature of face deepfakes in the wild, were developed. Also, newer variants of the pre-trained EfficientNET architecture such as EfficientNET, V2S, and V2M were used as the backbone of the face deepfake detection models in addition to the older variants seen in literature such as EfficientNET B3 and B5 to compare their performance against face deepfakes. More importantly, images generated by the latest StyleGAN generation technique called StyleGAN3 were included and applied in this work. The inclusion of StyleGAN3 generated images in a deep-fake detection model has not been observed in literature yet, therefore this paper can act as a starting point for reference on how pre-trained models such as EfficientNET perform against StyleGAN3. The results show the best accuracy rate is 92.77% achieved by the EfficientNET B3 backbone model, beating newer EfficientNET variants for the non-StyleGAN3 test set. However, when tested against a small pre-processed StyleGAN3 test set prepared separately, the best accuracy score achieved was only 64.29%. In addition, this was achieved by EfficientNET V2M which scored only 88.60% accuracy. The results demonstrate that even pretrained backbone-based detectors with very high performance against images generated by a variety of generation techniques including StyleGAN3 and some “in the wild” scenarios considered, struggle to retain their accuracy on a small set of pre-processed StyleGAN3 generated images.