<p>Multimodal sarcasm detection integrates textual, acoustic, and visual cues to identify sarcastic intent. However, a counter-intuitive pattern has emerged: state-of-the-art models often achieve better performance by excluding non-textual context, suggesting that naive integration introduces more interference than useful signal. We identify two factors underlying this degradation: modal noise from low-information content in acoustic and visual streams, and scenario noise from inherent ambiguity where identical features indicate opposite intents across different contexts. We propose MSNR-MSD, which addresses these challenges through complementary mechanisms. For modal noise, Dual-layer Modal Noise Reduction (DMNR) decomposes context refinement into redundancy filtering and target-driven fusion, enabling selective preservation of prosodic and emotional cues while removing interference. Empirical analysis reveals that redundancy filtering benefits acoustic but harms visual features, leading to modality-specific processing. For scenario noise, a lightweight calibration module leverages cross-sample prototypes to disambiguate high-uncertainty predictions. Experiments on MUStARD and MUStARD++ benchmarks demonstrate state-of-the-art performance, with ablation studies confirming that both mechanisms contribute consistent gains.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Modal and scenario noise reduction for multimodal sarcasm detection

  • Jianlin Chen,
  • Jiali Lin,
  • Dazhi Jiang

摘要

Multimodal sarcasm detection integrates textual, acoustic, and visual cues to identify sarcastic intent. However, a counter-intuitive pattern has emerged: state-of-the-art models often achieve better performance by excluding non-textual context, suggesting that naive integration introduces more interference than useful signal. We identify two factors underlying this degradation: modal noise from low-information content in acoustic and visual streams, and scenario noise from inherent ambiguity where identical features indicate opposite intents across different contexts. We propose MSNR-MSD, which addresses these challenges through complementary mechanisms. For modal noise, Dual-layer Modal Noise Reduction (DMNR) decomposes context refinement into redundancy filtering and target-driven fusion, enabling selective preservation of prosodic and emotional cues while removing interference. Empirical analysis reveals that redundancy filtering benefits acoustic but harms visual features, leading to modality-specific processing. For scenario noise, a lightweight calibration module leverages cross-sample prototypes to disambiguate high-uncertainty predictions. Experiments on MUStARD and MUStARD++ benchmarks demonstrate state-of-the-art performance, with ablation studies confirming that both mechanisms contribute consistent gains.