<p>Multimodal Emotion Recognition in Conversation (MERC) has emerged as a prominent research focus, aiming to accurately interpret emotional states by analyzing textual, audio, and visual modalities within dynamic conversational interactions. Despite significant progress in this field, existing methods still struggle with effectively modeling intra- and inter-modal dependencies, difficulty aligning heterogeneous feature representations, and capturing dynamic emotional transitions, which often leads to inconsistent predictions across dialogue turns. To overcome these limitations, we present a novel Graph-Guided Cross-Modal Attention Network (GGCMANet) equipped with Emotion Shift Modeling. First, we introduce a Multi-Graph Attention Mechanism (MGAM) that models each modality as a graph, where utterances are represented as nodes and both intra- and inter-modal dependencies are encoded as edges, thereby capturing intra-modal context and inter-modal features. Second, we propose a Cross-Modal Attention Network (CMAN) to align and integrate heterogeneous modality-specific features. By enabling dynamic cross-modal attention, the model selectively attends to the most informative inter-modal features, enhancing representation coherence and contextual alignment. Third, we design the Emotion Shift Model (ESM), which captures temporal variations in emotional states across dialogue turns and mitigates prediction inconsistencies over time. Extensive experiments on three benchmark datasets demonstrate that our approach consistently outperforms state-of-the-art methods in multimodal emotion recognition.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A graph-guided cross-modal attention network for multimodal emotion recognition via emotion-shift modeling

  • Sathiyamoorthi Arthanari,
  • Yeon-kug Moon

摘要

Multimodal Emotion Recognition in Conversation (MERC) has emerged as a prominent research focus, aiming to accurately interpret emotional states by analyzing textual, audio, and visual modalities within dynamic conversational interactions. Despite significant progress in this field, existing methods still struggle with effectively modeling intra- and inter-modal dependencies, difficulty aligning heterogeneous feature representations, and capturing dynamic emotional transitions, which often leads to inconsistent predictions across dialogue turns. To overcome these limitations, we present a novel Graph-Guided Cross-Modal Attention Network (GGCMANet) equipped with Emotion Shift Modeling. First, we introduce a Multi-Graph Attention Mechanism (MGAM) that models each modality as a graph, where utterances are represented as nodes and both intra- and inter-modal dependencies are encoded as edges, thereby capturing intra-modal context and inter-modal features. Second, we propose a Cross-Modal Attention Network (CMAN) to align and integrate heterogeneous modality-specific features. By enabling dynamic cross-modal attention, the model selectively attends to the most informative inter-modal features, enhancing representation coherence and contextual alignment. Third, we design the Emotion Shift Model (ESM), which captures temporal variations in emotional states across dialogue turns and mitigates prediction inconsistencies over time. Extensive experiments on three benchmark datasets demonstrate that our approach consistently outperforms state-of-the-art methods in multimodal emotion recognition.