<p>Speech emotion recognition is a critical research area in affective computing, with applications spanning human-computer interaction, virtual agent design, and emotion-driven decision-making systems. This study presents a novel speech emotion classification method. Initially, we employed the Geneva Wheel of Emotions (GWE) to categorize emotions into High Control (anger, happiness, disgust) and Low Control (fear, sadness, neutrality) classes, which align with the control dimension. The speech signals are then converted into log-Mel spectrograms, for which we propose DeepSpecCNN, a novel Convolutional Neural Network (CNN) architecture. The proposed model is further integrated into a strategic ensemble with other leading CNNs using soft voting, providing a powerful alternative to complex transformer-based structures. Experimental assessment on the Crowd-sourced Emotional Multimodal Actors Dataset (CREMA-D) demonstrates the efficacy of our approach. Our systematic investigation of various ensemble configurations reveals that a 5-model ensemble achieves a peak accuracy of <b>77.70%</b>, outperforming current state-of-the-art techniques by striking an ideal balance between classification sensitivity and computational efficiency.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Ensemble Learning for Improved Speech Emotion Recognition: A Control Dimension Analysis of Log-Mel Spectrograms from the CREMA-D Dataset

  • Zineddine Sarhani Kahhoul,
  • Nadjiba Terki,
  • Mohamed Lakhdar Tiar,
  • Ilyes Benaissa,
  • Selma Boutiba

摘要

Speech emotion recognition is a critical research area in affective computing, with applications spanning human-computer interaction, virtual agent design, and emotion-driven decision-making systems. This study presents a novel speech emotion classification method. Initially, we employed the Geneva Wheel of Emotions (GWE) to categorize emotions into High Control (anger, happiness, disgust) and Low Control (fear, sadness, neutrality) classes, which align with the control dimension. The speech signals are then converted into log-Mel spectrograms, for which we propose DeepSpecCNN, a novel Convolutional Neural Network (CNN) architecture. The proposed model is further integrated into a strategic ensemble with other leading CNNs using soft voting, providing a powerful alternative to complex transformer-based structures. Experimental assessment on the Crowd-sourced Emotional Multimodal Actors Dataset (CREMA-D) demonstrates the efficacy of our approach. Our systematic investigation of various ensemble configurations reveals that a 5-model ensemble achieves a peak accuracy of 77.70%, outperforming current state-of-the-art techniques by striking an ideal balance between classification sensitivity and computational efficiency.