<p>Speech Emotion Recognition (SER) aims to identify human emotions from vocal cues, enabling more natural and adaptive human–computer interaction. This study proposes a hybrid Deformable Convolutional Neural Network with Bidirectional Long Short-Term Memory (DCNN–BiLSTM) for robust SER using dual spectral features—Mel-Frequency Cepstral Coefficients (MFCCs) and Mel-spectrograms. The deformable convolutions provide adaptive spatial feature learning, while the BiLSTM layers capture temporal and contextual dependencies. Experiments were conducted on three benchmark corpora—RAVDESS, CREMA-D, and TESS—under speaker-independent and cross-corpus conditions. The proposed model achieved an overall accuracy of 82.4% and a macro-F1 score of 0.83, outperforming standard CNN and baseline DCNN models while maintaining computational efficiency. Enhanced regularization and class-balanced training eliminated prior prediction bias, ensuring stable multi-class performance. These results confirm the model’s suitability for real-time affective computing, mental-health monitoring, and intelligent virtual-assistant systems.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Learning emotional nuances in speech via DCNNs and spectral feature integration

  • K. Venkatesh Sharma,
  • Pramod Reddy Ayiluri,
  • Rakesh Betala,
  • Madhavi Pappula,
  • K. Shirisha Reddy

摘要

Speech Emotion Recognition (SER) aims to identify human emotions from vocal cues, enabling more natural and adaptive human–computer interaction. This study proposes a hybrid Deformable Convolutional Neural Network with Bidirectional Long Short-Term Memory (DCNN–BiLSTM) for robust SER using dual spectral features—Mel-Frequency Cepstral Coefficients (MFCCs) and Mel-spectrograms. The deformable convolutions provide adaptive spatial feature learning, while the BiLSTM layers capture temporal and contextual dependencies. Experiments were conducted on three benchmark corpora—RAVDESS, CREMA-D, and TESS—under speaker-independent and cross-corpus conditions. The proposed model achieved an overall accuracy of 82.4% and a macro-F1 score of 0.83, outperforming standard CNN and baseline DCNN models while maintaining computational efficiency. Enhanced regularization and class-balanced training eliminated prior prediction bias, ensuring stable multi-class performance. These results confirm the model’s suitability for real-time affective computing, mental-health monitoring, and intelligent virtual-assistant systems.