<p>Facial expression recognition (FER) represents a crucial component in modern human-computer interaction, affective computing, and intelligent surveillance applications. Despite significant advances in deep learning approaches, existing methods are fundamentally limited by label ambiguity in unconstrained “in-the-wild” environments, where crowd-sourced annotations often mismatch the inherently uncertain and mixed nature of human emotions. This mismatch causes deep models to overfit noisy supervision instead of learning robust expression patterns, resulting in poor generalization performance. In response to these challenges, we propose an expression aware supervision and refinement (EASR) framework that incorporates a novel expression aware loss (EAL). The EAL incorporates a learnable expression similarity structure that dynamically penalizes confusions between semantically close expressions more softly while maintaining strong penalties for distant misclassifications, effectively reshaping the training signal into cost-sensitive, ambiguity-aware supervision. This mechanism enables standard FER architectures trained on single hard labels to benefit from structured, soft-label-like guidance without requiring multi-annotator distributions or additional labeling efforts. Comprehensive experimental results on four challenging in-the-wild datasets (RAF-DB, FERPlus, AffectNet, and CEFE) demonstrate that EASR achieves substantial performance improvements, with accuracies of 90.16%, 90.66%, 65.60%, and 68.33%, respectively, reaching or surpassing recent state-of-the-art methods under comparable model complexity. Our code is available at <a href="https://github.com/fr36/EASR">https://github.com/fr36/EASR</a></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Easr: expression aware supervision and refinement for in-the-wild facial expression recognition

  • Xuantao Nie,
  • Zixiang Fei,
  • Yukun Zhang,
  • Wenju Zhou,
  • Minrui Fei,
  • Xia Li

摘要

Facial expression recognition (FER) represents a crucial component in modern human-computer interaction, affective computing, and intelligent surveillance applications. Despite significant advances in deep learning approaches, existing methods are fundamentally limited by label ambiguity in unconstrained “in-the-wild” environments, where crowd-sourced annotations often mismatch the inherently uncertain and mixed nature of human emotions. This mismatch causes deep models to overfit noisy supervision instead of learning robust expression patterns, resulting in poor generalization performance. In response to these challenges, we propose an expression aware supervision and refinement (EASR) framework that incorporates a novel expression aware loss (EAL). The EAL incorporates a learnable expression similarity structure that dynamically penalizes confusions between semantically close expressions more softly while maintaining strong penalties for distant misclassifications, effectively reshaping the training signal into cost-sensitive, ambiguity-aware supervision. This mechanism enables standard FER architectures trained on single hard labels to benefit from structured, soft-label-like guidance without requiring multi-annotator distributions or additional labeling efforts. Comprehensive experimental results on four challenging in-the-wild datasets (RAF-DB, FERPlus, AffectNet, and CEFE) demonstrate that EASR achieves substantial performance improvements, with accuracies of 90.16%, 90.66%, 65.60%, and 68.33%, respectively, reaching or surpassing recent state-of-the-art methods under comparable model complexity. Our code is available at https://github.com/fr36/EASR