The correct distinction between highly realistic computer-generated (CG) images and photographic (PG) images has become an important area of research. In recent years, most of the CG image forensics methods are proposed based on deep learning, but the detection performances of these methods still need to be improved, especially in terms of robustness and generalization. To tackle these issues, we leverage the Vision Transformer (ViT) model, which excels in capturing the global features of images, and design a Forensic Feature Pre-processing (FFP) module to further improve the detection performance. Experiments are conducted on a large-scale CG image benchmark (LSCGB), which is a challenging dataset for CG image detection. The proposed approach can achieve high detection accuracy. Extensive experiments on different public datasets and common post-processing operations demonstrate our approach can achieve significantly better generalization and robustness than the state-of-the-art approaches.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Computer-Generated Image Forensics Based on Vision Transformer with Forensic Feature Pre-processing Module

  • Yifang Chen,
  • Guanchen Wen,
  • Yong Wang,
  • Jianhua Yang,
  • Yu Zhang

摘要

The correct distinction between highly realistic computer-generated (CG) images and photographic (PG) images has become an important area of research. In recent years, most of the CG image forensics methods are proposed based on deep learning, but the detection performances of these methods still need to be improved, especially in terms of robustness and generalization. To tackle these issues, we leverage the Vision Transformer (ViT) model, which excels in capturing the global features of images, and design a Forensic Feature Pre-processing (FFP) module to further improve the detection performance. Experiments are conducted on a large-scale CG image benchmark (LSCGB), which is a challenging dataset for CG image detection. The proposed approach can achieve high detection accuracy. Extensive experiments on different public datasets and common post-processing operations demonstrate our approach can achieve significantly better generalization and robustness than the state-of-the-art approaches.