RELUEM-Reinforcing Emotional Understanding: Advancing Speech Emotion Recognition Through Deep Reinforcement Learning
摘要
Speech emotion recognition (SER) is a critical component in domains such as human–computer interaction, mental health monitoring, and affective computing. While existing approaches—particularly those based on deep learning—have made significant progress, they often struggle to accurately capture subtle emotional cues in speech, especially in real-world scenarios. A key gap remains in the adaptability and decision-making capabilities of current models when faced with dynamic and ambiguous emotional expressions. To address this, we propose a novel speech emotion recognition framework that integrates deep reinforcement learning (DRL) with conventional classification models. Unlike static models, our approach enables adaptive learning and optimal decision policies for emotion classification over time. We evaluate our method using the widely adopted RAVDESS benchmark dataset. Experimental results demonstrate a notable improvement in performance, with our method achieving an accuracy of 89.5% and an F1 score of 0.86—outperforming several state-of-the-art baselines. These findings suggest that DRL offers a promising pathway toward more robust and context-aware speech emotion recognition systems.