Emotion Pattern Recognition of Speech Signals Using Representation Learning with Limited Annotations
摘要
Accurate identification of emotions from speech signals is an open area of research. Discriminating features, as well as, annotations are hard to obtain and identify the emotional state and its intensity in valence-arousal scale. In this work, we propose a multistage model for emotion recognition using speech signal. Proposed method exploits the strength of representation learning to learn a compact latent representation of the speech signal. This captures the prominent markers in two-dimensional valence and arousal space. Further, we perform hierarchical clustering using the learned latent representation exploiting a best choice of clustering and distance measure. Our method during learning, inherently identifies the clusters in two stages - arousal and valence, respectively. We have experimented using publicly available speech emotion dataset. We perform the inferencing in arousal-level and valence-arousal scale with average accuracy of 88.5% and 82.3% respectively across multiple users, which outperforms existing state-of-the-art methods.