<p>Emotion recognition from speech is a significant research area in human–computer interaction and psychological assessments. This study proposes a novel three-stage process for emotion recognition from speech signals. In the first stage, pre-processing operations, including noise reduction and normalization, are applied to improve data quality. Then, StarGAN is employed for data augmentation, addressing the challenge of limited emotional speech data. The second stage involves extracting meaningful features from high-dimensional data using deep convolutional neural networks (DCNN). Finally, support vector machines (SVM) are used to classify the extracted features into distinct emotional categories. The main advantage of the proposed method over traditional deep learning techniques is the use of StarGAN for data augmentation, which improves performance on limited datasets. Additionally, the combination of DCNN and SVM enhances accuracy and efficiency in feature extraction and classification. This method has been evaluated using four different datasets, RAVDESS, Emo-DB, SAVEE, and CASIA, achieving accuracy rates of 98.25%, 95.5%, 96.8%, and 96.45%, respectively. The results indicate that the proposed method significantly outperforms existing methods and offers improved capabilities for practical application. Despite achieving high accuracy, challenges such as increased computational complexity and variability in performance across different datasets are discussed.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-Model emotion recognition from speech using StarGAN, DCNN, and SVM

  • Najmeh Sadat Banihosseini,
  • Vahid Ghods

摘要

Emotion recognition from speech is a significant research area in human–computer interaction and psychological assessments. This study proposes a novel three-stage process for emotion recognition from speech signals. In the first stage, pre-processing operations, including noise reduction and normalization, are applied to improve data quality. Then, StarGAN is employed for data augmentation, addressing the challenge of limited emotional speech data. The second stage involves extracting meaningful features from high-dimensional data using deep convolutional neural networks (DCNN). Finally, support vector machines (SVM) are used to classify the extracted features into distinct emotional categories. The main advantage of the proposed method over traditional deep learning techniques is the use of StarGAN for data augmentation, which improves performance on limited datasets. Additionally, the combination of DCNN and SVM enhances accuracy and efficiency in feature extraction and classification. This method has been evaluated using four different datasets, RAVDESS, Emo-DB, SAVEE, and CASIA, achieving accuracy rates of 98.25%, 95.5%, 96.8%, and 96.45%, respectively. The results indicate that the proposed method significantly outperforms existing methods and offers improved capabilities for practical application. Despite achieving high accuracy, challenges such as increased computational complexity and variability in performance across different datasets are discussed.