Conversational Emotion Detection (CED), spanning across multiple modalities (e.g., textual, visual and acoustic modalities), has been drawing ever-more interest in the multi-modal fields. Previous studies consistently consider the CED task as an emotion classification problem utterance by utterance, which largely ignore the global topic information of each conversation, especially the multi-modal topic information inside multiple modalities. Obviously, such information is crucial for alleviating the emotional information deficiency problem in a single utterance. With this in mind, we propose a Topic-enriched Variational Transformer (TVT) approach to capture the conversational topic information inside different modalities for CED. Particularly, a modality-independent topic module in TVT is designed to mine topic clues from either the discrete textual content, or the continuous visual and acoustic contents in each conversation. Detailed evaluation shows the great advantage of TVT to the CED task over the state-of-the-art baselines, justifying the importance of the multi-modal topic information to CED and the effectiveness of our approach in capturing such information.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Topic-Enriched Variational Transformer for Conversational Emotion Detection

  • Jiamin Luo,
  • Jingjing Wang,
  • Guodong Zhou

摘要

Conversational Emotion Detection (CED), spanning across multiple modalities (e.g., textual, visual and acoustic modalities), has been drawing ever-more interest in the multi-modal fields. Previous studies consistently consider the CED task as an emotion classification problem utterance by utterance, which largely ignore the global topic information of each conversation, especially the multi-modal topic information inside multiple modalities. Obviously, such information is crucial for alleviating the emotional information deficiency problem in a single utterance. With this in mind, we propose a Topic-enriched Variational Transformer (TVT) approach to capture the conversational topic information inside different modalities for CED. Particularly, a modality-independent topic module in TVT is designed to mine topic clues from either the discrete textual content, or the continuous visual and acoustic contents in each conversation. Detailed evaluation shows the great advantage of TVT to the CED task over the state-of-the-art baselines, justifying the importance of the multi-modal topic information to CED and the effectiveness of our approach in capturing such information.