Recent achievements in deep learning have provided a thrust in the development of emotion recognition systems. Researchers have developed multimodal emotion recognition models that leverage both facial and non-facial data such as text and audio. However, most of the models suffer from the well-known problem of imbalanced classes which occurs due to the unavailability of data in many emotion categories. Motivated by this, the current article has proposed an imbalance-aware multimodal attention-enabled latent space oversampling framework to address the uneven distribution of data samples in various emotion categories. The primary contributions of the manuscript are as follows: (i) A multimodal deep learning framework is developed to predict emotion from facial images and textual data. (ii) The model includes a separate attention-enabled CNN-based image encoder to extract a latent representation of input image. (iii) A separate BERT encoder is used to extract textual features. (iv) Both latent vectors are fused to obtain a multimodal latent representation of both image and text data. Thereafter, the latent space oversampling method is used to address the class imbalance problem. The balanced latent vectors are then employed to train shallow learning models to predict the emotion. Experiments have revealed that the attention-enabled latent space oversampling framework can effectively predict emotion from multimodal data with greater accuracy than multimodal models without attention and latent space oversampling.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Imbalance-Aware Multimodal Attention-Enabled Latent Space Oversampling for Emotion Recognition

  • Sanchita Das,
  • Ritoja Mukhopadhyay,
  • Prerana Dutta,
  • Diya Soor,
  • Arya Sarkar,
  • Sankhadeep Chatterjee,
  • Asit Kumar Das

摘要

Recent achievements in deep learning have provided a thrust in the development of emotion recognition systems. Researchers have developed multimodal emotion recognition models that leverage both facial and non-facial data such as text and audio. However, most of the models suffer from the well-known problem of imbalanced classes which occurs due to the unavailability of data in many emotion categories. Motivated by this, the current article has proposed an imbalance-aware multimodal attention-enabled latent space oversampling framework to address the uneven distribution of data samples in various emotion categories. The primary contributions of the manuscript are as follows: (i) A multimodal deep learning framework is developed to predict emotion from facial images and textual data. (ii) The model includes a separate attention-enabled CNN-based image encoder to extract a latent representation of input image. (iii) A separate BERT encoder is used to extract textual features. (iv) Both latent vectors are fused to obtain a multimodal latent representation of both image and text data. Thereafter, the latent space oversampling method is used to address the class imbalance problem. The balanced latent vectors are then employed to train shallow learning models to predict the emotion. Experiments have revealed that the attention-enabled latent space oversampling framework can effectively predict emotion from multimodal data with greater accuracy than multimodal models without attention and latent space oversampling.