Learning emotional nuances in speech via DCNNs and spectral feature integration
摘要
Speech Emotion Recognition (SER) aims to identify human emotions from vocal cues, enabling more natural and adaptive human–computer interaction. This study proposes a hybrid Deformable Convolutional Neural Network with Bidirectional Long Short-Term Memory (DCNN–BiLSTM) for robust SER using dual spectral features—Mel-Frequency Cepstral Coefficients (MFCCs) and Mel-spectrograms. The deformable convolutions provide adaptive spatial feature learning, while the BiLSTM layers capture temporal and contextual dependencies. Experiments were conducted on three benchmark corpora—RAVDESS, CREMA-D, and TESS—under speaker-independent and cross-corpus conditions. The proposed model achieved an overall accuracy of 82.4% and a macro-F1 score of 0.83, outperforming standard CNN and baseline DCNN models while maintaining computational efficiency. Enhanced regularization and class-balanced training eliminated prior prediction bias, ensuring stable multi-class performance. These results confirm the model’s suitability for real-time affective computing, mental-health monitoring, and intelligent virtual-assistant systems.