MAC-DAG: a modality-adaptive contextual directed acyclic graph network for multimodal emotion recognition in conversations
摘要
Multimodal conversational emotion recognition aims to accurately infer emotional states in dialogues by integrating textual, acoustic, and visual information. However, existing approaches often overlook the contextual discrepancies across different modalities, making it challenging to simultaneously capture semantic, temporal, and speaker-level dependencies. To address these issues, we propose a Modality-Adaptive Contextual Directed Acyclic Graph Network (MAC-DAG). Specifically, MAC-DAG constructs modality-specific directed acyclic graphs–semantic-driven, temporal-driven, and speaker-driven–to enable differentiated contextual modeling and causal propagation within each modality, and employs a Dual-Channel Context Propagation (DCP) mechanism to facilitate dynamic interaction between nodes and their contextual representations. For cross-modal fusion, a Cross-modal Coupled Attention (CCA) mechanism is introduced, which incorporates an Early-stage Semantic Alignment and Text-guided Enhancement (ECSA) module and a Deep Cross-modal Interaction and Gated Fusion (DCGF) module to achieve dynamic alignment and complementary integration across modalities. In addition, a Self-distillation Feedback Mechanism (SDFM) is adopted to jointly optimize contextual graph modeling and multimodal fusion-based classification, ensuring structural consistency and semantic coordination. Experimental results on the IEMOCAP and MELD datasets show that the MAC-DAG model achieves promising performance, which verifies its certain effectiveness in multimodal emotion modeling within complex conversational scenarios.