<p>Facial micro-expressions (ME) are fast, minute movements of the facial muscles that communicate emotions and intentions that are concealed. Due to limited datasets, ephemeral occurrence, subtle movement of facial muscles, subjective classification, noise and background, and ethical quandaries, it is difficult to identify ME from videos. This study proposes the deep model DITRAGS-STNet and the FHOG-DITRAGS-STNet comprising the dynamic texture handcrafted feature descriptor HOG-TOP and the optical flow feature descriptor Bi-WOOF, and the smaller size DITRAGS-STNet deep model to classify ME videos. Both models use deep channels that integrate dense CNN blocks, customized transition blocks with factorized and residual inception, and GRUs with soft temporal attention to create high-level features. The random forest ensemble classifier is used to categorize the feature vectors into three classes: negative, positive, and surprise. FHOG-DITRAGS-STNet fuses the complementary capabilities of three feature sources to improve classification performance. Two composite datasets are constructed: MASPOWER, which mitigates overfitting through domain adaptation, and CASSMEW-MICRO, which enhances ME recognition across diverse samples while increasing sample size. These datasets are created using five benchmark macro-expression datasets and six benchmark ME datasets, respectively. Extensive cross-validations demonstrate that both models are superior in terms of efficacy and generality. </p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Facial micro-expression recognition from videos through domain adaptation and multi-modal spatio-temporal feature ensemble

  • MD. Sajjatul Islam,
  • Yongsheng Sang,
  • Adam A. Q. Mohammed,
  • Jiancheng Lv

摘要

Facial micro-expressions (ME) are fast, minute movements of the facial muscles that communicate emotions and intentions that are concealed. Due to limited datasets, ephemeral occurrence, subtle movement of facial muscles, subjective classification, noise and background, and ethical quandaries, it is difficult to identify ME from videos. This study proposes the deep model DITRAGS-STNet and the FHOG-DITRAGS-STNet comprising the dynamic texture handcrafted feature descriptor HOG-TOP and the optical flow feature descriptor Bi-WOOF, and the smaller size DITRAGS-STNet deep model to classify ME videos. Both models use deep channels that integrate dense CNN blocks, customized transition blocks with factorized and residual inception, and GRUs with soft temporal attention to create high-level features. The random forest ensemble classifier is used to categorize the feature vectors into three classes: negative, positive, and surprise. FHOG-DITRAGS-STNet fuses the complementary capabilities of three feature sources to improve classification performance. Two composite datasets are constructed: MASPOWER, which mitigates overfitting through domain adaptation, and CASSMEW-MICRO, which enhances ME recognition across diverse samples while increasing sample size. These datasets are created using five benchmark macro-expression datasets and six benchmark ME datasets, respectively. Extensive cross-validations demonstrate that both models are superior in terms of efficacy and generality.