Multimodal Test-Time Adaptation for Fake News Detection
摘要
Most existing approaches to fake news detection focus on leveraging multimodal semantic information through fusion or interaction. However, they rarely explore the relationships and differences in social media data, often assuming that the training and testing datasets share the same distribution. To address those challenges, we introduce T3FND, a novel framework that leverages Test-Time Training (TTT) for multimodal fake news detection. Our approach integrates a Masked Autoencoder (MAE) to enhance the model’s ability to capture nuanced relationships and discrepancies within multimodal data. Additionally, we propose a Test-Time Training strategy designed to improve model generalization by utilizing the fine-grained features learned by the MAE from the test data. To tackle the issue of mismatched image-text pairs in fake news datasets during TTT, we develop the Multi-modal Mask-Transformer (M3-Transformer) module. Extensive experiments conducted on two widely-used datasets convincingly demonstrate that our method can greatly improve the performance.