<p>Facial expression recognition, as an important task in the field of computer vision, infers a person’s emotion by analyzing facial expression characteristics. With the continuous development of deep learning, great progress has been made in facial expression recognition, and in order to further improve the accuracy and robustness of facial expression recognition, this paper proposes a novel recognition network based on multi-scale feature fusion and enhanced squeeze-excitation global attention. It is found that despite the high diversity and complexity of facial expressions, existing techniques have limitations in extracting important facial features. Therefore, in this paper, spatial-channel convolution strategies are introduced to expand the receptive field of the network to capture the global information of the image more efficiently. In addition, to integrate feature information at different scales and enhance the attention to important regions of facial expressions, this paper designs a feature fusion strategy and an enhanced squeeze-excitation global attention strategy. Extensive experiments on FER2013, CK+, JAFFE, FERPlus, and AffectNet datasets demonstrate the superiority of our method, which achieves competitive performance with accuracy of 74.39%, 98.23%, 97.14%, 84.83%, and 66.83% respectively, surpassing several existing approaches by a considerable margin.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-scale feature fusion and enhanced squeeze-excitation global attention based facial expression recognition

  • Dongdong Wang,
  • Songhao Zhu,
  • Zhiwei Liang

摘要

Facial expression recognition, as an important task in the field of computer vision, infers a person’s emotion by analyzing facial expression characteristics. With the continuous development of deep learning, great progress has been made in facial expression recognition, and in order to further improve the accuracy and robustness of facial expression recognition, this paper proposes a novel recognition network based on multi-scale feature fusion and enhanced squeeze-excitation global attention. It is found that despite the high diversity and complexity of facial expressions, existing techniques have limitations in extracting important facial features. Therefore, in this paper, spatial-channel convolution strategies are introduced to expand the receptive field of the network to capture the global information of the image more efficiently. In addition, to integrate feature information at different scales and enhance the attention to important regions of facial expressions, this paper designs a feature fusion strategy and an enhanced squeeze-excitation global attention strategy. Extensive experiments on FER2013, CK+, JAFFE, FERPlus, and AffectNet datasets demonstrate the superiority of our method, which achieves competitive performance with accuracy of 74.39%, 98.23%, 97.14%, 84.83%, and 66.83% respectively, surpassing several existing approaches by a considerable margin.