Accurate identification of emotions from speech signals is an open area of research. Discriminating features, as well as, annotations are hard to obtain and identify the emotional state and its intensity in valence-arousal scale. In this work, we propose a multistage model for emotion recognition using speech signal. Proposed method exploits the strength of representation learning to learn a compact latent representation of the speech signal. This captures the prominent markers in two-dimensional valence and arousal space. Further, we perform hierarchical clustering using the learned latent representation exploiting a best choice of clustering and distance measure. Our method during learning, inherently identifies the clusters in two stages - arousal and valence, respectively. We have experimented using publicly available speech emotion dataset. We perform the inferencing in arousal-level and valence-arousal scale with average accuracy of 88.5% and 82.3% respectively across multiple users, which outperforms existing state-of-the-art methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Emotion Pattern Recognition of Speech Signals Using Representation Learning with Limited Annotations

  • Anish Datta,
  • Soma Bandyopadhyay,
  • Gauri Deshpande,
  • Ramesh Ramakrishnan,
  • Arpan Pal

摘要

Accurate identification of emotions from speech signals is an open area of research. Discriminating features, as well as, annotations are hard to obtain and identify the emotional state and its intensity in valence-arousal scale. In this work, we propose a multistage model for emotion recognition using speech signal. Proposed method exploits the strength of representation learning to learn a compact latent representation of the speech signal. This captures the prominent markers in two-dimensional valence and arousal space. Further, we perform hierarchical clustering using the learned latent representation exploiting a best choice of clustering and distance measure. Our method during learning, inherently identifies the clusters in two stages - arousal and valence, respectively. We have experimented using publicly available speech emotion dataset. We perform the inferencing in arousal-level and valence-arousal scale with average accuracy of 88.5% and 82.3% respectively across multiple users, which outperforms existing state-of-the-art methods.