<p>The rapid evolution of AI-generated synthetic media, called deepfakes, has raised substantial concerns regarding digital misinformation, security breaches, and public trust. Existing centralized detection systems are often limited with privacy risks, weak generalizability, and reduced robustness when exposed to multimodal manipulations across audio, text, and video data. There is always a human cognitive capacity to fuse and contrast cues across sensory modalities while judging authenticity. This research proposes a privacy-preserving, federated learning-based deep-fake detection framework that facilitates secure, decentralized training across heterogeneous devices. The proposed framework leverages Fed-DFakeNet for localized model training, enabling robust feature extraction from multimodal datasets without transmitting raw data. It incorporates MogDetNet, a knowledge distillation-based fusion module that aligns multimodal features for improved generalization and accuracy. Furthermore, the mCreamFL aggregation strategy introduces a contrastive representation ensemble and encrypted communication mechanism, ensuring data integrity and optimal performance while preserving user privacy. Comprehensive experiments are conducted using three benchmark datasets FoR (audio), SDFVD (video), and TweepFake (text) demonstrating superior performance. The framework achieves classification accuracies of 98.32%, 99.43%, and 99.87%, significantly outperforming state-of-the-art baselines. Evaluation metrics such as sensitivity, F-measure, and G-mean reinforce the framework’s reliability, robustness, and resilience against adversarial content. This study contributes a scalable and secure deepfake detection pipeline and underscores the critical importance of privacy-aware AI systems in combating the increasing sophistication of generative media. The results affirm that combining multimodal feature learning with federated intelligence offers a powerful and ethically aligned solution to emerging digital threats.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cognitive-Inspired Multimodal Deepfake Detection Using Privacy-Preserving Federated Intelligence

  • Aisha Blfgeh,
  • Hanin Ardah

摘要

The rapid evolution of AI-generated synthetic media, called deepfakes, has raised substantial concerns regarding digital misinformation, security breaches, and public trust. Existing centralized detection systems are often limited with privacy risks, weak generalizability, and reduced robustness when exposed to multimodal manipulations across audio, text, and video data. There is always a human cognitive capacity to fuse and contrast cues across sensory modalities while judging authenticity. This research proposes a privacy-preserving, federated learning-based deep-fake detection framework that facilitates secure, decentralized training across heterogeneous devices. The proposed framework leverages Fed-DFakeNet for localized model training, enabling robust feature extraction from multimodal datasets without transmitting raw data. It incorporates MogDetNet, a knowledge distillation-based fusion module that aligns multimodal features for improved generalization and accuracy. Furthermore, the mCreamFL aggregation strategy introduces a contrastive representation ensemble and encrypted communication mechanism, ensuring data integrity and optimal performance while preserving user privacy. Comprehensive experiments are conducted using three benchmark datasets FoR (audio), SDFVD (video), and TweepFake (text) demonstrating superior performance. The framework achieves classification accuracies of 98.32%, 99.43%, and 99.87%, significantly outperforming state-of-the-art baselines. Evaluation metrics such as sensitivity, F-measure, and G-mean reinforce the framework’s reliability, robustness, and resilience against adversarial content. This study contributes a scalable and secure deepfake detection pipeline and underscores the critical importance of privacy-aware AI systems in combating the increasing sophistication of generative media. The results affirm that combining multimodal feature learning with federated intelligence offers a powerful and ethically aligned solution to emerging digital threats.