<p>With the advancement of information technology, video data have become a primary source of information on the internet, shaping public perception. However, maliciously tampered videos pose a serious threat to social trust, highlighting the necessity for video authentication. This study focuses on the detection of video frame deletion manipulation. Although existing frame deletion forensic methods have achieved remarkable progress, most rely solely on visual data, limiting their performances in complex scenarios. To address this limitation, we propose a novel framework, TMTC (Trustworthy Multimodal Transformer Classification), which integrates both audio and visual features for improved detection performance. Specifically, the framework leverages an enhanced ResNet to extract audio features and a three-dimensional convolutional neural network (3DCNN) to capture visual features. The multimodal fusion method, which is based on uncertainty, may yield conflicts in confidence levels during the feature fusion process. To facilitate effective multimodal fusion, we propose a cross-modal feature mediation (CFM) module that addresses modality-specific confidence bias and resolves inter-modal discrepancies, including temporal misalignment and feature inconsistency. Finally, a Dempster-Shafer evidence theory module is utilized for robust classification. Experimental results show that TMTC achieves a 1.48% accuracy improvement on non-degraded datasets and a 5.43% improvement on noisy datasets, compared to the best-performing method among the state-of-the-art frame deletion detection techniques evaluated.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TMTC: trusted multi-modal transformer classification framework for video frame deletion detection

  • Chunhui Feng,
  • Yongxiang Zhong,
  • Yigong Huang,
  • Xiaolong Liu

摘要

With the advancement of information technology, video data have become a primary source of information on the internet, shaping public perception. However, maliciously tampered videos pose a serious threat to social trust, highlighting the necessity for video authentication. This study focuses on the detection of video frame deletion manipulation. Although existing frame deletion forensic methods have achieved remarkable progress, most rely solely on visual data, limiting their performances in complex scenarios. To address this limitation, we propose a novel framework, TMTC (Trustworthy Multimodal Transformer Classification), which integrates both audio and visual features for improved detection performance. Specifically, the framework leverages an enhanced ResNet to extract audio features and a three-dimensional convolutional neural network (3DCNN) to capture visual features. The multimodal fusion method, which is based on uncertainty, may yield conflicts in confidence levels during the feature fusion process. To facilitate effective multimodal fusion, we propose a cross-modal feature mediation (CFM) module that addresses modality-specific confidence bias and resolves inter-modal discrepancies, including temporal misalignment and feature inconsistency. Finally, a Dempster-Shafer evidence theory module is utilized for robust classification. Experimental results show that TMTC achieves a 1.48% accuracy improvement on non-degraded datasets and a 5.43% improvement on noisy datasets, compared to the best-performing method among the state-of-the-art frame deletion detection techniques evaluated.