Emotion detection in Urdu speech: a deep hybrid learning approach
摘要
Speech Emotion Recognition (SER) aims to identify human emotions from vocal expressions, contributing significantly to affective computing applications. Despite progress in SER for widely spoken languages, Urdu remains underrepresented due to the lack of dedicated datasets and language-specific modeling. This study proposes a deep learning framework specifically designed for Urdu SER, utilizing robust acoustic features including Mel Frequency Cepstral Coefficients (MFCC) with delta and double-delta derivatives, as well as Mel-spectrograms, to capture both spectral and temporal aspects of speech. We test and analyze three different model architectures: a standalone Convolutional Neural Network (CNN) with attention mechanism, a CNN-LSTM hybrid, and a CNN-LSTM hybrid with an attention mechanism. The attention-based CNN-LSTM model had the highest accuracy 95%, efficiently capturing subtle emotional cues in Urdu speech. These results highlight the potential of attention mechanisms and hybrid architectures in advancing emotion recognition for under-resourced languages like Urdu.