Micro-expression (ME) is a spontaneous psychological stress response characterized by subtle and fleeting facial movements, which can reflect an individual’s genuine emotion. Recognizing ME accurately can provide crucial technical support for fields such as lie detection, hence attracting increasing attention from researchers in psychology and artificial intelligence. However, current automated ME recognition (MER) methods typically require ME data captured by high-speed cameras to achieve good performance, which is challenging to perform effectively in the wild with low frame-rate. To this end, we propose a novel cross frame-rate representation alignment MER framework, termed H2LMER, to enhance the low frame-rate ME feature learning. Specifically, we first customize a progressive training strategy and a cross frame-rate similarity loss to align the representation of high and low frame-rate ME data in emotion feature space. Additionally, to eliminate the interference of facial identity, we introduce a triplet loss for discriminative emotion feature learning. Afterwards, we transfer the effectively pretrained feature extractor to target low frame-rate ME datasets by partial layers finetuning. Finally, extensive experiments on CAS(ME) \(^3\) PART A and C datasets have varified the effectiveness of our H2LMER, which outperforms single-stage baselines and existing state-of-the-art MER methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

H2LMER: A Cross Frame-Rate Representation Alignment Framework for Micro-expression Recognition

  • Xinglong Mao,
  • Shifeng Liu,
  • Sirui Zhao,
  • Yiming Zhang,
  • Hao Wang,
  • Tong Xu,
  • Enhong Chen

摘要

Micro-expression (ME) is a spontaneous psychological stress response characterized by subtle and fleeting facial movements, which can reflect an individual’s genuine emotion. Recognizing ME accurately can provide crucial technical support for fields such as lie detection, hence attracting increasing attention from researchers in psychology and artificial intelligence. However, current automated ME recognition (MER) methods typically require ME data captured by high-speed cameras to achieve good performance, which is challenging to perform effectively in the wild with low frame-rate. To this end, we propose a novel cross frame-rate representation alignment MER framework, termed H2LMER, to enhance the low frame-rate ME feature learning. Specifically, we first customize a progressive training strategy and a cross frame-rate similarity loss to align the representation of high and low frame-rate ME data in emotion feature space. Additionally, to eliminate the interference of facial identity, we introduce a triplet loss for discriminative emotion feature learning. Afterwards, we transfer the effectively pretrained feature extractor to target low frame-rate ME datasets by partial layers finetuning. Finally, extensive experiments on CAS(ME) \(^3\) PART A and C datasets have varified the effectiveness of our H2LMER, which outperforms single-stage baselines and existing state-of-the-art MER methods.