Facial Expression Recognition (FER) in unconstrained environments is often hindered by challenges such as pose variations and occlusions, which degrade model performance and robustness. To address these limitations, we propose EMO-Net, a lightweight and robust FER framework that incorporates dual-branch feature fusion and channel-wise attention mechanisms. The core innovation of EMO-Net is its Feature Fusion Module (FFM), which integrates semantically rich and geometrically precise features through a dual-branch architecture, preserving their complementary properties and reducing redundancy. This integration enhances the model’s ability to maintain facial context awareness, improving its understanding of the entire face. Additionally, a novel Local Attention Module (LAM) leverages channel-wise attention mechanisms to focus on critical facial regions, capturing fine-grained expression details while minimizing computational cost. Extensive experiments on five widely-used FER benchmark datasets demonstrate that EMO-Net outperforms existing methods, achieving state-of-the-art performance, particularly under challenging conditions involving occlusions and pose variations.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

EMO-Net: A Lightweight Framework for Robust Facial Expression Recognition via Dual-Branch Fusion and Channel-Wise Attention

  • Lingyan Zhang,
  • Qing Zhong,
  • Yu Xia,
  • Li Kuang,
  • Yiman Xie

摘要

Facial Expression Recognition (FER) in unconstrained environments is often hindered by challenges such as pose variations and occlusions, which degrade model performance and robustness. To address these limitations, we propose EMO-Net, a lightweight and robust FER framework that incorporates dual-branch feature fusion and channel-wise attention mechanisms. The core innovation of EMO-Net is its Feature Fusion Module (FFM), which integrates semantically rich and geometrically precise features through a dual-branch architecture, preserving their complementary properties and reducing redundancy. This integration enhances the model’s ability to maintain facial context awareness, improving its understanding of the entire face. Additionally, a novel Local Attention Module (LAM) leverages channel-wise attention mechanisms to focus on critical facial regions, capturing fine-grained expression details while minimizing computational cost. Extensive experiments on five widely-used FER benchmark datasets demonstrate that EMO-Net outperforms existing methods, achieving state-of-the-art performance, particularly under challenging conditions involving occlusions and pose variations.