Stacking Ensemble Learning with 1D CNN and CapsuleNet for Speech Emotion Recognition
摘要
Emotion intelligence is playing a crucial role in identifying the mental and physical health of humans. Corporates, government organizations, and schools are concentrating on mental health for better working environment. Automization of identifying emotions will enable the process of measuring the emotional intelligence. Speech Emotion Recognition (SER) is the best way to identify human emotions. In this paper, MFCC features from the non-linguistic details of speech are extracted for training the Stacking Ensemble Learner. Advanced neural networks, viz., 1D CNN and CapsuleNets are stacked as base learners, and multilayer perceptron is used at metalayer for decision-making. The experiments are conducted on two benchmark datasets RAVDESS and IEMOCAP. The classification accuracies achieved are 91.57% and 91.81%, respectively. Precision, recall, and F1-scores are also evaluated, and the results obtained reflect the proficiency of the proposed model over the existing model. Also, the classification of happy emotion, which is very crucial in SER, has encouraging results.