DTDiff: adaptive decoupled transformer with language-conditioned denoising learning for multimodal emotion recognition in conversation
摘要
Multimodal Emotion Recognition in Conversation (MERC) has attracted significant attention in recent years, and existing methods mainly rely on contextual cues and multimodal interactions to predict emotions. However, these methods often suffer from the detrimental effects of noise, including contextual interaction noise and modality-specific noise contamination, leading to suboptimal model performance. Therefore, we propose DTDiff, a noise-aware framework that tackles two types of noise: inter-utterance interaction noise through an Adaptive Residual Decoupled Transformer (ARDT), and modality-specific noise via Language-Conditioned Denoising Learning (LDL). Specifically, ARDT improves robustness by effectively filtering irrelevant contextual dependencies and enhances the representation of each modality through decoupled residual attention fusion. Meanwhile, LDL employs language-conditioned diffusion models to denoise visual and acoustic modalities and measures noise levels via gating mechanisms. Finally, we introduce Dual-Signal Alignment to further promote multimodal fusion. Experiments on IEMOCAP and MELD datasets demonstrate that DTDiff outperforms state-of-the-art methods.