Speech Emotion Recognition Using Convolutional Neural Networks
摘要
An efficient conversation requires understanding of speech along with accurate interpretation of the tone and emotion of the speaker. Misinterpretation of emotions is one of the biggest communication barriers. It affects personal communication, business dealings, customer handling, interviews, etc. It occurs because of the simple fact that people think, speak and interpret differently. Speech Emotion Recognition, often abbreviated as SER, is based on the observation that emotions are frequently conveyed through voice, in the form of tone and pitch variations. With the rise of service industry, exact interpretation and correctly gauging a message has become a basic requirement. Multiple industries can leverage this capability to provide a range of services. For instance, marketing companies can recommend products based on a person’s emotions, call center representatives can interact more effectively by adapting to the customer’s mood, and the auto-motive sector can detect a person’s emotions to adjust autonomous vehicle speed for collision avoidance. Consequently, this type of application holds significant potential in the world, offering benefits to companies and enhancing consumer safety. The motivation behind this paper was to build a machine learning model capable of discerning emotions conveyed through speech, assessing and comprehending an individual’s emotional state, and subsequently providing personalized recommendations based on their mood in the future. The proposed Spectrogram-MFCC fusion model’s impressive accuracy of 97% in detecting emotions represents a significant advancement, particularly in applications such as chatbots and social robots, where the ability to perceive concealed emotions in speech is vital for fostering meaningful interactions.