Optimization of Speech Emotion Recognition Based on Deep Learning Techniques
摘要
Speech Emotion Recognition (SER) is essential for applications such as emergency situations, virtual reality encounters, interaction between humans and robots, and behavior analysis. Conventional methods extract significant characteristics from voice spectrograms using hand-crafted features and CNN, which frequently leads to high computing complexity. Using a key sequence segment selection technique and radial basis function network similarity measurement in the produced clusters, we present a novel SER framework. The discovered critical segments are transformed into spectrograms using the Short-Time Fourier Transform (STFT), which are further processed by a CNN to extract discriminative features. After being normalized to ensure proper detection, these features are put into a deep bi-directional long short-term memory network to gain the temporal information required for ultimate emotion recognition. Our strategy improves recognition accuracy by processing critical portions rather than the complete phrase. Experiments on the RAVDESS datasets demonstrate the suggested model’s efficiency and reliability, which surpasses the most sophisticated techniques with an accuracy of 84.15%.