<p>Emotion detection plays a crucial role in human computer interactions, as it enhances the user experience by ensuring that systems respond appropriately user’s emotions. However, recognizing emotions in videos is challenging due to factors such as facial expressions, body language, background, occlusions, and dynamic object positions, all of which interfere with the interpretation of facial and limbs movements. To address these challenges, we presented an approach that incorporated dynamic key frame selection, YOLOv5 object recognition and the Augmented-Attention Convolutional Transformer Fusion Network (AACTFN). The dynamic key frame selection algorithm identifies the top five key frame pairs with significant differences by calculating contextual mean loss between Deepflow and ResNet18 to generate their top five deep flow images. These deep flow images are then processed by YOLOv5, which identifies emotion-dense elements, such as faces and bodies, while discarding background noise. The resulting feature maps are processed through AACTFN, which combines the spatial learning of CNNs with the temporal modeling of transformers. As a result, our approach effectively captures complex emotional patterns, even in dynamic occlusion scenarios. By employing deep spatio-temporal analysis for feature extraction, our method achieves state-of-the-art accuracy and reliability in emotion detection on the CAER, CK+, DFEW, FERV39k, and RAVDESS datasets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Video-based emotion recognition using motion-aware deep hybrid learning

  • Navneet Gupta,
  • R. Vishnu Priya,
  • Chandan Kumar Verma

摘要

Emotion detection plays a crucial role in human computer interactions, as it enhances the user experience by ensuring that systems respond appropriately user’s emotions. However, recognizing emotions in videos is challenging due to factors such as facial expressions, body language, background, occlusions, and dynamic object positions, all of which interfere with the interpretation of facial and limbs movements. To address these challenges, we presented an approach that incorporated dynamic key frame selection, YOLOv5 object recognition and the Augmented-Attention Convolutional Transformer Fusion Network (AACTFN). The dynamic key frame selection algorithm identifies the top five key frame pairs with significant differences by calculating contextual mean loss between Deepflow and ResNet18 to generate their top five deep flow images. These deep flow images are then processed by YOLOv5, which identifies emotion-dense elements, such as faces and bodies, while discarding background noise. The resulting feature maps are processed through AACTFN, which combines the spatial learning of CNNs with the temporal modeling of transformers. As a result, our approach effectively captures complex emotional patterns, even in dynamic occlusion scenarios. By employing deep spatio-temporal analysis for feature extraction, our method achieves state-of-the-art accuracy and reliability in emotion detection on the CAER, CK+, DFEW, FERV39k, and RAVDESS datasets.