Deep Learning-Based Speech Emotion Recognition with Reference to Gender Separation
摘要
Humans have always found language to be the most basic and immediate way to communicate with one another. The logical advancement is to extend this kind of communication to realm of computers. A range of techniques are employed by speech emotion recognition systems (SER) to identify spoken emotional expressions. This field of research that explores subjective emotions that speakers communicate has been described as speech emotion recognition. Tone and pitch are powerful emotional indicators, and our method makes use of that. This paper tries to recognize emotion in a short voice message. This paper utilizes the TESS and SAVEE datasets, which encompass seven primary emotions: happiness, fear, anger, disgust, surprise, sadness, or neutrality. Input voice samples represent different emotions in.wav files and use MFCC feature extraction to eliminate unnecessary noise. Subsequently, organize collected sequential data into a 3D array suitable for the CNN machine learning (ML) model. Results from the experiments show that the suggested model has an average accuracy of 89.07%.