<p>Facial action units (AUs) serve as a precise descriptor of facial expressions, revealing an individual’s psychological and mental state. Due to the fact that each AU is confined to a specific facial region, AU feature extraction usually necessitates the integration of landmark detection tools and prior knowledge regarding to locations of different AUs to partition the face, which leads to time-consuming and laborsome pre-processing procedure. To tackle this issue, a weakly supervised guided attention inference network is proposed for AU detection. The network encompasses two modules with shared parameters: a classification and ROI segmentation module (CRSM) and an attention mining module (AMM). The CRSM autonomously identifies regions of interest for target AUs and generates class activation attention maps. The AMM utilizes these maps to exclude facial regions of target AUs so that a weak constraint that minimizes AU prediction scores for target AUs can be imposed on network training, which thereby ensures that the network’s attention maps encompass all most discriminant regions contributing to AU classification decisions. Experimental results on the BP4D and DISFA datasets demonstrate that, even in the absence of landmark detection and pre-facial region partitioning, the proposed model sustains excellent detection performance during testing. Furthermore, the generated AU attention map can accurately indicate the spatial locations of AU occurrences, which makes the AU detection results explainable.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Facial AU detection based on a guided attention inference network with embedded regional segmentation branch

  • Kui Li,
  • Chaolei Liang,
  • Wei Zou,
  • Danfeng Hu,
  • JiaJun Wang

摘要

Facial action units (AUs) serve as a precise descriptor of facial expressions, revealing an individual’s psychological and mental state. Due to the fact that each AU is confined to a specific facial region, AU feature extraction usually necessitates the integration of landmark detection tools and prior knowledge regarding to locations of different AUs to partition the face, which leads to time-consuming and laborsome pre-processing procedure. To tackle this issue, a weakly supervised guided attention inference network is proposed for AU detection. The network encompasses two modules with shared parameters: a classification and ROI segmentation module (CRSM) and an attention mining module (AMM). The CRSM autonomously identifies regions of interest for target AUs and generates class activation attention maps. The AMM utilizes these maps to exclude facial regions of target AUs so that a weak constraint that minimizes AU prediction scores for target AUs can be imposed on network training, which thereby ensures that the network’s attention maps encompass all most discriminant regions contributing to AU classification decisions. Experimental results on the BP4D and DISFA datasets demonstrate that, even in the absence of landmark detection and pre-facial region partitioning, the proposed model sustains excellent detection performance during testing. Furthermore, the generated AU attention map can accurately indicate the spatial locations of AU occurrences, which makes the AU detection results explainable.