<p>Facial Emotion Recognition (FER) plays a crucial role in understanding human behavior and interactions, with applications in human-computer interaction, healthcare, surveillance, and multimedia content analysis. Emotion recognition from speech videos is particularly challenging as it requires capturing and analyzing subtle variations in facial expressions along with dynamic speech changes. In this paper, a hybrid learning approach is presented that integrates machine learning and deep learning models to enhance the accuracy and robustness of visual facial emotion recognition in speech videos. The proposed approach leverages a pre-trained ResNet-101 model for feature extraction and evaluates three widely used machine learning classifiers, Random Decision Forest (RDF), Decision Tree Boosting (DTB), and Support Vector Machine (SVM) for emotion prediction. Experiments conducted on the RAVDESS dataset demonstrate that our hybrid model, combining Deep CNN with RDF classification, achieves a high accuracy of 92.50%, outperforming existing state-of-the-art models. This work highlights the potential of integrating deep learning and machine learning for FER, offering significant benefits for applications in healthcare, education, and human-computer interaction.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hybrid Learning Based Visual Facial Emotion Recognition in Speech Videos

  • Yogesh Rochlani,
  • A. B. Raut

摘要

Facial Emotion Recognition (FER) plays a crucial role in understanding human behavior and interactions, with applications in human-computer interaction, healthcare, surveillance, and multimedia content analysis. Emotion recognition from speech videos is particularly challenging as it requires capturing and analyzing subtle variations in facial expressions along with dynamic speech changes. In this paper, a hybrid learning approach is presented that integrates machine learning and deep learning models to enhance the accuracy and robustness of visual facial emotion recognition in speech videos. The proposed approach leverages a pre-trained ResNet-101 model for feature extraction and evaluates three widely used machine learning classifiers, Random Decision Forest (RDF), Decision Tree Boosting (DTB), and Support Vector Machine (SVM) for emotion prediction. Experiments conducted on the RAVDESS dataset demonstrate that our hybrid model, combining Deep CNN with RDF classification, achieves a high accuracy of 92.50%, outperforming existing state-of-the-art models. This work highlights the potential of integrating deep learning and machine learning for FER, offering significant benefits for applications in healthcare, education, and human-computer interaction.