Multimodal Emotion Recognition in Conversation (MERC) utilizes multimodal information such as language, visual, and audio to enhance the understanding of human emotions. Current multimodal interaction frameworks inadequately resolve inherent information conflicts and redundancy due to their assumption of equivalent quality across heterogeneous modalities. In addition, inappropriate evaluation of the importance of modalities can also cause this problem. To address this issue, we introduce a Language-Focused Augmented Transformer with Variational Distillation Fusion network called LFVD. In contrast to previous work, we suggest focusing on language modality through the Language-Focused Augmented Transformer, which extracts task-relevant signals from visual and audio modalities to help us understand language. Concurrently, this architecture derives conversational emotional atmosphere representation to refine multimodal integration, thereby mitigating the influence of redundant and conflicting information. Furthermore, Variational Distillation Fusion has been proposed in which multimodal representations are probabilistically encoded as variational distributions over Gaussian manifolds rather than deterministic embeddings. Subsequently, the importance of each modality is estimated automatically based on distribution differences. Experiments conducted on the IEMOCAP and MELD datasets indicate that our proposed model has better performance than previous state-of-the-art baseline models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LFVD: Language-Focused Augmented Transformer with Variational Distillation Fusion for Multimodal Emotion Recognition in Conversation

  • Shuming Jiang,
  • Tingting Zhang,
  • Xiaofei Zhu

摘要

Multimodal Emotion Recognition in Conversation (MERC) utilizes multimodal information such as language, visual, and audio to enhance the understanding of human emotions. Current multimodal interaction frameworks inadequately resolve inherent information conflicts and redundancy due to their assumption of equivalent quality across heterogeneous modalities. In addition, inappropriate evaluation of the importance of modalities can also cause this problem. To address this issue, we introduce a Language-Focused Augmented Transformer with Variational Distillation Fusion network called LFVD. In contrast to previous work, we suggest focusing on language modality through the Language-Focused Augmented Transformer, which extracts task-relevant signals from visual and audio modalities to help us understand language. Concurrently, this architecture derives conversational emotional atmosphere representation to refine multimodal integration, thereby mitigating the influence of redundant and conflicting information. Furthermore, Variational Distillation Fusion has been proposed in which multimodal representations are probabilistically encoded as variational distributions over Gaussian manifolds rather than deterministic embeddings. Subsequently, the importance of each modality is estimated automatically based on distribution differences. Experiments conducted on the IEMOCAP and MELD datasets indicate that our proposed model has better performance than previous state-of-the-art baseline models.