Emotion recognition is the task of identifying the emotion of each corpus in a conversation, and the task of multimodal emotion recognition (MER) in multi-party conversations is more challenging compared to the traditional emotion recognition, and most of the existing work is based on Transfomer for inter-modal information fusion. In this paper, we propose Multimodal Mamba Model for Emotion Recognition (MMME), a Mamba-based model for multimodal emotion recognition, which captures inter-modal interactions through cross-modal Mamba in order to learn better modal representations, and the computational complexity of Mamba’s approximate linearity can be somewhat lower compared to that of the Attention mechanism can improve the training and inference speed of the model and reduce the number of parameters of the model to some extent. Experiments on the publicly available English emotion recognition datasets IEMOCAP, MELD show that MMME outperforms previous state-of-the-art baselines.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multimodal Mamba Model for Emotion Recognition in Conversations

  • TaoZheng Zhang,
  • Zhaoyang Chen,
  • Jiantao Du

摘要

Emotion recognition is the task of identifying the emotion of each corpus in a conversation, and the task of multimodal emotion recognition (MER) in multi-party conversations is more challenging compared to the traditional emotion recognition, and most of the existing work is based on Transfomer for inter-modal information fusion. In this paper, we propose Multimodal Mamba Model for Emotion Recognition (MMME), a Mamba-based model for multimodal emotion recognition, which captures inter-modal interactions through cross-modal Mamba in order to learn better modal representations, and the computational complexity of Mamba’s approximate linearity can be somewhat lower compared to that of the Attention mechanism can improve the training and inference speed of the model and reduce the number of parameters of the model to some extent. Experiments on the publicly available English emotion recognition datasets IEMOCAP, MELD show that MMME outperforms previous state-of-the-art baselines.