<p>In recent years, speech emotion recognition has attracted significant interest in the area of affective computing, human-computer interaction, and mental health evaluation. This paper introduces a strategy to enhance emotion recognition by converting speech signals into visual representations called chaograms, which are then processed using deep convolutional neural networks (DCNNs). In this method, the speech signal is first represented in a 3D reconstructed phase space (RPS). The parameters of the RPS have been set according to the speaker’s gender, which is determined through a former gender classification stage. Then, by calculating the projections of the data points formed in this space along the three coordinate axes, three 2D tensors are generated. These tensors are treated as the red, green, and blue (RGB) channels of an image to form the chaogram. Chaograms can effectively capture the dynamic and temporal structure of speech signals, enabling more robust feature extraction. The proposed method employs DCNNs such as AlexNet, VGG16, InceptionV3, and ResNet50 to recognize and classify various emotions. By leveraging the power of deep learning, our approach reduces reliance on handcrafted features, leading to improved generalization across different datasets. Transfer learning and fine-tuning techniques are applied to the DCNNs to enhance accuracy rates and speed up the training process. Additionally, data augmentation techniques and hyperparameter optimization are employed to improve the robustness of the presented deep learning models against variations in speech recordings. The proposed models were tested on the two publicly available datasets, EMO-DB and eNTERFACE05, and the results prove that the suggested approach outperforms previously published methods. The employed AlexNet, VGG16, InceptionV3, and ResNet50 models provide accuracy rates of 93.54%, 97.16%, 97.56%, and 97.68% on the EMO-DB dataset and 89.14%, 91.45%, 92.08%, and 92.34% on the eNTERFACE05 dataset.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An approach to accurate recognition of emotions through speech-to-image signal conversion and deep convolutional neural networks

  • Mohammad Reza Falahzadeh,
  • Yazdan ZandiyeVakili,
  • Ali Harimi,
  • Edris Zaman Farsa,
  • Arash Ahmadi,
  • Ajith Abraham

摘要

In recent years, speech emotion recognition has attracted significant interest in the area of affective computing, human-computer interaction, and mental health evaluation. This paper introduces a strategy to enhance emotion recognition by converting speech signals into visual representations called chaograms, which are then processed using deep convolutional neural networks (DCNNs). In this method, the speech signal is first represented in a 3D reconstructed phase space (RPS). The parameters of the RPS have been set according to the speaker’s gender, which is determined through a former gender classification stage. Then, by calculating the projections of the data points formed in this space along the three coordinate axes, three 2D tensors are generated. These tensors are treated as the red, green, and blue (RGB) channels of an image to form the chaogram. Chaograms can effectively capture the dynamic and temporal structure of speech signals, enabling more robust feature extraction. The proposed method employs DCNNs such as AlexNet, VGG16, InceptionV3, and ResNet50 to recognize and classify various emotions. By leveraging the power of deep learning, our approach reduces reliance on handcrafted features, leading to improved generalization across different datasets. Transfer learning and fine-tuning techniques are applied to the DCNNs to enhance accuracy rates and speed up the training process. Additionally, data augmentation techniques and hyperparameter optimization are employed to improve the robustness of the presented deep learning models against variations in speech recordings. The proposed models were tested on the two publicly available datasets, EMO-DB and eNTERFACE05, and the results prove that the suggested approach outperforms previously published methods. The employed AlexNet, VGG16, InceptionV3, and ResNet50 models provide accuracy rates of 93.54%, 97.16%, 97.56%, and 97.68% on the EMO-DB dataset and 89.14%, 91.45%, 92.08%, and 92.34% on the eNTERFACE05 dataset.