ERF-BA-TFD+: a multimodal model for audio-visual deepfake detection
摘要
Deepfake detection is a critical task in identifying manipulated multimedia content. In real-world scenarios, deepfake content can manifest across multiple modalities, including audio and video. To address this challenge, we propose a novel multimodal deepfake detection model, ERF-BA-TFD+, which combines an enhanced receptive field (ERF) and audio-visual fusion. Our model processes both audio and video features simultaneously, leveraging their complementary information to improve detection accuracy and robustness. The key innovation of ERF-BA-TFD+ lies in its ability to model long-range dependencies within the audio-visual input, allowing it to better capture subtle discrepancies between real and fake content. In our experiments, we evaluate ERF-BA-TFD+ on the LAV-DF and DDL-AV datasets, which consist of segmented and full-length video clips. Unlike previous benchmarks that primarily focused on isolated segments, these datasets enable us to assess the model’s performance in a more comprehensive and realistic setting. Our method achieves state-of-the-art results on the DDL-AV dataset and outperforms most models on the LAV-DF dataset, demonstrating superior accuracy and processing speed compared to existing techniques. Additionally, the ERF-BA-TFD+ model showcased its effectiveness in the “Workshop on Deepfake Detection, Localization, and Interpretability,” Track 2: Audio-Visual Detection and Localization (DDL-AV), and won the first place in this competition.