Facial Expression Recognition (FER) has been a major development in a number of sectors in recent years. To reduce the problems caused by numerous parameters and high computational resource requirements of the current FER model, we propose a lightweight FER method based on Hybrid Multiscale and Multi-Head Collaborative Attention. Firstly, the EfficientNet_b0 is further streamlined by adjusting the number of channels. Then, to better extract multiscale features, we design the Hybrid Multiscale module to improve its capacity to extract features. Secondly, a Multi-Head Collaborative Attention module is intended to direct the model's attention to multiple facial key areas. Finally, the outputs of different attention-heads are fused in the Regional Attention module. A Regional Attention loss is used to force the model to focus on diverse regions to avoid focusing on the same area repeatedly. The experiments are conducted on publicly accessible real-scene datasets RAF-DB, FERPlus, AffectNet7, and AffectNet8. The proposed method achieves recognition accuracies of 88.49%, 89.40%, 64.09%, and 60.52%, respectively. The number of parameters is only 2.1M, and it uses fewer parameters to obtain a performance that is nearly identical to that of the most advanced techniques now in use.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Lightweight Facial Expression Recognition Based on Hybrid Multiscale and Multi-Head Collaborative Attention

  • Haitao Zhang,
  • Xufei Zhuang,
  • Xudong Gao,
  • Rui Mao,
  • Qing-Dao-Er-Ji Ren

摘要

Facial Expression Recognition (FER) has been a major development in a number of sectors in recent years. To reduce the problems caused by numerous parameters and high computational resource requirements of the current FER model, we propose a lightweight FER method based on Hybrid Multiscale and Multi-Head Collaborative Attention. Firstly, the EfficientNet_b0 is further streamlined by adjusting the number of channels. Then, to better extract multiscale features, we design the Hybrid Multiscale module to improve its capacity to extract features. Secondly, a Multi-Head Collaborative Attention module is intended to direct the model's attention to multiple facial key areas. Finally, the outputs of different attention-heads are fused in the Regional Attention module. A Regional Attention loss is used to force the model to focus on diverse regions to avoid focusing on the same area repeatedly. The experiments are conducted on publicly accessible real-scene datasets RAF-DB, FERPlus, AffectNet7, and AffectNet8. The proposed method achieves recognition accuracies of 88.49%, 89.40%, 64.09%, and 60.52%, respectively. The number of parameters is only 2.1M, and it uses fewer parameters to obtain a performance that is nearly identical to that of the most advanced techniques now in use.