<p>Facial expression recognition (FER) is a fundamental task in affective computing, enabling machines to understand and respond to human emotions in domains such as mental health assessment, human–robot interaction, and intelligent tutoring systems. This study presents FER-HA, a hybrid attention-based model that integrates high-level visual features extracted from EfficientNet-B3 with geometric descriptors derived from 68 facial landmarks. An adaptive attention mechanism dynamically weights the multimodal representations to emphasize emotion-relevant information and mitigate challenges such as subtle expression variations and partial occlusions. Evaluated on the FER-2013 benchmark, FER-HA achieves a test accuracy of 72.32%, averaged over three independent runs, representing a 4.2% improvement over the baseline EfficientNet-B3. Notably, the model shows substantial gains for minority emotion classes such as disgust (F1 = 0.73) and fear (F1 = 0.53). With only 12.1 million parameters, the proposed model achieves a strong balance between accuracy and computational efficiency, making it suitable for real-time and resource-constrained systems. These results highlight the effectiveness of attention-guided multimodal fusion for scalable and efficient emotion recognition in supercomputing environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

FER-HA: a hybrid attention model for facial emotion recognition

  • Reza Nemati,
  • Kimia Shirini,
  • Sina Samadi Gharehveran

摘要

Facial expression recognition (FER) is a fundamental task in affective computing, enabling machines to understand and respond to human emotions in domains such as mental health assessment, human–robot interaction, and intelligent tutoring systems. This study presents FER-HA, a hybrid attention-based model that integrates high-level visual features extracted from EfficientNet-B3 with geometric descriptors derived from 68 facial landmarks. An adaptive attention mechanism dynamically weights the multimodal representations to emphasize emotion-relevant information and mitigate challenges such as subtle expression variations and partial occlusions. Evaluated on the FER-2013 benchmark, FER-HA achieves a test accuracy of 72.32%, averaged over three independent runs, representing a 4.2% improvement over the baseline EfficientNet-B3. Notably, the model shows substantial gains for minority emotion classes such as disgust (F1 = 0.73) and fear (F1 = 0.53). With only 12.1 million parameters, the proposed model achieves a strong balance between accuracy and computational efficiency, making it suitable for real-time and resource-constrained systems. These results highlight the effectiveness of attention-guided multimodal fusion for scalable and efficient emotion recognition in supercomputing environments.