Topic-Enriched Variational Transformer for Conversational Emotion Detection
摘要
Conversational Emotion Detection (CED), spanning across multiple modalities (e.g., textual, visual and acoustic modalities), has been drawing ever-more interest in the multi-modal fields. Previous studies consistently consider the CED task as an emotion classification problem utterance by utterance, which largely ignore the global topic information of each conversation, especially the multi-modal topic information inside multiple modalities. Obviously, such information is crucial for alleviating the emotional information deficiency problem in a single utterance. With this in mind, we propose a Topic-enriched Variational Transformer (TVT) approach to capture the conversational topic information inside different modalities for CED. Particularly, a modality-independent topic module in TVT is designed to mine topic clues from either the discrete textual content, or the continuous visual and acoustic contents in each conversation. Detailed evaluation shows the great advantage of TVT to the CED task over the state-of-the-art baselines, justifying the importance of the multi-modal topic information to CED and the effectiveness of our approach in capturing such information.