Joint Multi-cue Learning for Emotion Recognition in Human-Computer Interaction
摘要
In this study, we propose a novel emotion recognition model that addresses the challenges of hierarchical labels and incorporates both face and body cues for joint learning in the context of human-computer interaction. We design the Graph Attention Body module (GAB), which utilizes the topology graph and attention mechanisms to extract features from body cues for joint learning with the extracted features from face cues. For the optimization approach of joint learning, we present the Loss Gradient Optimization module (LGO) that leverages multi-label multi-loss features to optimize the loss calculation. The results demonstrate that our method outperforms existing approaches in terms of Accuracy and \(F1-score\) . Our findings highlight the significance of multi-cue joint learning for emotion recognition in human-computer interaction.