<p>Emotion recognition in conversations, as a key technology for affective dialogue systems, requires efficient computational performance in real-time applications to ensure a smooth user experience. Existing research primarily focuses on interlocutor modelling and conversational information extraction, often neglecting the simultaneous effective capture of both short-range and long-range contextual semantic information in conversations, and fails to fully consider complementary information and interaction dynamics between modalities, which affects the accuracy of emotion recognition and overall system performance. To address these issues, this paper proposes the MEA&amp;CMDF model, aiming to maximize emotion recognition accuracy while fully leveraging high-performance computing resources. First, this model integrates the fine-grained interaction modelling capability of single-layer multi-head attention mechanisms in local contexts along with their inherent parallel computing advantages, combined with the long-range dependency modelling capability of the Mamba state space model and its linear complexity characteristics, thereby constructing a hybrid module to effectively capture both short-range and long-range contextual semantic information in dialogues. Second, to address the limitations of multimodal feature fusion, a dynamic interactive fusion network is designed that precisely locates key emotional cues in each modality through cross-modal attention mechanisms, and utilizes graph neural networks to model complex cross-modal dependencies, achieving adaptive weight allocation and deep complementary fusion of modal information. Extensive experiments on two widely used benchmark datasets, IEMOCAP and MELD, demonstrate that MEA&amp;CMDF achieves significant breakthroughs in capturing short- and long-range emotional dependencies and synergistic utilization of multimodal information, effectively improving the accuracy and robustness of multimodal conversational emotion recognition, providing a feasible solution for high-performance computing in real-time dialogue systems.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Mamba-based dynamic fusion of cross-modal attention optimization models for ERC

  • Jun Wu,
  • Yu Chen,
  • Panpan Chen,
  • Shuai Guo,
  • Jiahui Huang,
  • Xinyi Zhu,
  • Qun Zhang

摘要

Emotion recognition in conversations, as a key technology for affective dialogue systems, requires efficient computational performance in real-time applications to ensure a smooth user experience. Existing research primarily focuses on interlocutor modelling and conversational information extraction, often neglecting the simultaneous effective capture of both short-range and long-range contextual semantic information in conversations, and fails to fully consider complementary information and interaction dynamics between modalities, which affects the accuracy of emotion recognition and overall system performance. To address these issues, this paper proposes the MEA&CMDF model, aiming to maximize emotion recognition accuracy while fully leveraging high-performance computing resources. First, this model integrates the fine-grained interaction modelling capability of single-layer multi-head attention mechanisms in local contexts along with their inherent parallel computing advantages, combined with the long-range dependency modelling capability of the Mamba state space model and its linear complexity characteristics, thereby constructing a hybrid module to effectively capture both short-range and long-range contextual semantic information in dialogues. Second, to address the limitations of multimodal feature fusion, a dynamic interactive fusion network is designed that precisely locates key emotional cues in each modality through cross-modal attention mechanisms, and utilizes graph neural networks to model complex cross-modal dependencies, achieving adaptive weight allocation and deep complementary fusion of modal information. Extensive experiments on two widely used benchmark datasets, IEMOCAP and MELD, demonstrate that MEA&CMDF achieves significant breakthroughs in capturing short- and long-range emotional dependencies and synergistic utilization of multimodal information, effectively improving the accuracy and robustness of multimodal conversational emotion recognition, providing a feasible solution for high-performance computing in real-time dialogue systems.