<p>Current human-computer dialogue often appears rigid and lacks empathy, making multimodal emotion recognition in conversations (MERC) essential for improving interaction quality. MERC faces three major challenges: (1) fully leveraging contextual information in dialogue; (2) capturing the nonlinear relationships among features from different modalities; and (3) distinguishing between semantically similar but categorically different few-shot emotion labels. To address these challenges, we propose the following strategies. First, we design the CoMamba framework, which extracts emotional cues from the dialogue context of each modality in depth, considering both global and local semantic perspectives. Second, we propose KAN-Fuse to replace traditional linear networks, aiming to capture the complex nonlinear relationships among emotional features. Moreover, this network facilitates more accurate modeling of few-shot emotional features. Extensive experiments conducted on two benchmark datasets for ERC demonstrate that our model outperforms state-of-the-art approaches, with particularly notable improvements in recognizing rare and semantically similar emotions.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CMKF:Multimodal emotion recognition in conversations based on CoMamba and KAN-Fuse

  • Junxiang Min,
  • Xianying Huang,
  • Yu Cheng

摘要

Current human-computer dialogue often appears rigid and lacks empathy, making multimodal emotion recognition in conversations (MERC) essential for improving interaction quality. MERC faces three major challenges: (1) fully leveraging contextual information in dialogue; (2) capturing the nonlinear relationships among features from different modalities; and (3) distinguishing between semantically similar but categorically different few-shot emotion labels. To address these challenges, we propose the following strategies. First, we design the CoMamba framework, which extracts emotional cues from the dialogue context of each modality in depth, considering both global and local semantic perspectives. Second, we propose KAN-Fuse to replace traditional linear networks, aiming to capture the complex nonlinear relationships among emotional features. Moreover, this network facilitates more accurate modeling of few-shot emotional features. Extensive experiments conducted on two benchmark datasets for ERC demonstrate that our model outperforms state-of-the-art approaches, with particularly notable improvements in recognizing rare and semantically similar emotions.