Bangla Speech Emotion Recognition (BSER) is an approach where a computer is able to identify human emotions from Bangla speech and enhance human-computer interaction. In this study, we propose an innovative framework for Bangla SER by integrating data augmentation, advanced feature extraction, and a deep learning-based classification model. We utilised three publicly available Bangla speech datasets: BanglaSER, KBES, and SUBESCO. To enhance the quantity and diversity of our dataset, we employed data augmentation techniques such as noise injection and time shifting. Previous Bangla Speech Emotion Recognition (SER) systems had high misclassification rates, particularly in lower tonal classes, suggesting that we require more robust and discriminative feature extraction approaches. To do this, we gathered different sound features such as Short-Time Fourier Transform (STFT), Chroma STFT, Mel Spectrogram, and Mel-Frequency Cepstral Coefficients (MFCC) to help the model classify emotions better. In this research, we use a deep learning model named 3-Dimensional Convolutional Neural Network (3D-CNN) to identify emotional states from Bangla speech datasets and have achieved an accuracy of 98.26%. Due to the limited availability of Bangla speech datasets and to assess the model’s resilience, we perform cross-validation with lingual datasets, namely, TESS, RAVDEES, EmoDB, and EMOVO. Impressively, our model shows high accuracy rates of 100%, 95%, 94%, and 94% on these datasets, respectively.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Bangla Speech Emotion Recognition Using 3D-CNN: A Multi-corpus and Cross-Lingual Study

  • Md Habibul Islam,
  • Nitun Kumar Podder,
  • Md. Nazmus Sakib,
  • Ankur Kumar Mondol,
  • Md Raihanul Haque,
  • Mst. Shaima Aslam Chaity

摘要

Bangla Speech Emotion Recognition (BSER) is an approach where a computer is able to identify human emotions from Bangla speech and enhance human-computer interaction. In this study, we propose an innovative framework for Bangla SER by integrating data augmentation, advanced feature extraction, and a deep learning-based classification model. We utilised three publicly available Bangla speech datasets: BanglaSER, KBES, and SUBESCO. To enhance the quantity and diversity of our dataset, we employed data augmentation techniques such as noise injection and time shifting. Previous Bangla Speech Emotion Recognition (SER) systems had high misclassification rates, particularly in lower tonal classes, suggesting that we require more robust and discriminative feature extraction approaches. To do this, we gathered different sound features such as Short-Time Fourier Transform (STFT), Chroma STFT, Mel Spectrogram, and Mel-Frequency Cepstral Coefficients (MFCC) to help the model classify emotions better. In this research, we use a deep learning model named 3-Dimensional Convolutional Neural Network (3D-CNN) to identify emotional states from Bangla speech datasets and have achieved an accuracy of 98.26%. Due to the limited availability of Bangla speech datasets and to assess the model’s resilience, we perform cross-validation with lingual datasets, namely, TESS, RAVDEES, EmoDB, and EMOVO. Impressively, our model shows high accuracy rates of 100%, 95%, 94%, and 94% on these datasets, respectively.