Decoupled Representation with Multimodal Prompts for Emotion Recognition in Conversation
摘要
Emotion recognition in conversation (ERC) has attracted significant attention due to its wide application in human-computer interaction. Information from various forms, such as text, audio, and video, can collectively enhance and complement the analysis of emotional context in conversations. However, there are still certain limitations in exploring emotional relationships in dialogues. Most existing methods focus on investigating correlations across different modalities or designing complex fusion strategies. However, due to the distribution gaps between modalities, directly fusing multimodal information often results in unrefined features. Therefore, in this work, we propose a novel model for Decoupled Representation with Multimodal Prompts (DRMP). First, we design the modality-shared encoder and the modality-specific encoder to obtain modality-shared and modality-specific features. Furthermore, since most existing methods overlook the inherent priors in modality information, we design text prompts for both types of features to further optimize the modality representations. Finally, we adopt a transformer-based model to capture both intra- and inter-modality interactions. Experimental results demonstrate that our method is competitive.