ST-SBV: Spatial-Temporal Self-Blended Videos for Deepfake Detection
摘要
The generalization ability to unseen forgery data is a critical concern in deepfake detection tasks. In this paper, we propose a novel spatial-temporal generation method for synthesizing forgery data, called spatial-temporal self-blended video (ST-SBV). ST-SBV is specifically designed to train deepfake detectors in a self-supervised manner. We utilize various image augmentation techniques on a genuine video to create pseudo source and target videos. Subsequently, ST-SBV is generated by blending them using face masks, following the generation process of a deepfake video. The key idea behind ST-SBV is to imitate the spatial and temporal artifacts in deepfake videos that are inherent and generalizable across deepfake generation methods. This encourages detectors to learn more general and effective representations. Specially, compared to existing data synthesis methods, ST-SBV addresses the gap in investigating spatial-temporal inconsistencies in forgery videos. Extensive experiments demonstrate that our method significantly improves the generalization performance of deepfake detectors on unknown forgery data. In particular, on the challenging datasets DFDCP and DFDC, our method outperforms the baselines by \(5.20\%\) and \(4.94\%\) , respectively.