Micro expressions (MEs), revealing a person’s genuine psychological activities, have extensive applications in areas such as law enforcement and mental health. MEs are characterized by their subtle amplitude and brief duration, which often confuse with macro expressions (MaEs). In this case, spotting such quick and subtle movements from long videos poses a significant challenge. Nevertheless, traditional deep models face obstacles in capturing the fine-grained features of MEs due to the insufficient samples. To address this issue, a three-stream neural network, named three-stream Convolutional Transformer fusion model (TCTFM), is proposed to strengthen diverse learning representation on MEs spotting. Specifically, this study takes two adjacent frames from the video, processed with crop-align technique, to calculate optical flow features and emphasize specific facial areas by the Regions of Interest selection. These optical flow features are then fed to the three-stream convolution module for feature extraction. Subsequently, the extracted features are fed into the Swin-Transformer block and cross attention module for both intra and inter features interaction. Finally, MEs spotting on TCTFM seeks to predict a score indicating the probability that a frame falls within a certain expression range. The proposed spotting approach demonstrated exceptional performance in the MEGC 2021 spotting task, achieving the overall F1 score of 0.3547 and 0.3976 on the CAS(ME)2 dataset and the SAMM LV dataset, respectively. Experimental results show that MEs spotting based on TCTFM boost the representation capability on informative MEs details and outperforms existing MEs spotting models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Three Streams Convolutional Transformer Fusion Model for Facial Macro- and Micro-expressions Spotting

  • Zhihua Xie,
  • Zhiwu Zhou

摘要

Micro expressions (MEs), revealing a person’s genuine psychological activities, have extensive applications in areas such as law enforcement and mental health. MEs are characterized by their subtle amplitude and brief duration, which often confuse with macro expressions (MaEs). In this case, spotting such quick and subtle movements from long videos poses a significant challenge. Nevertheless, traditional deep models face obstacles in capturing the fine-grained features of MEs due to the insufficient samples. To address this issue, a three-stream neural network, named three-stream Convolutional Transformer fusion model (TCTFM), is proposed to strengthen diverse learning representation on MEs spotting. Specifically, this study takes two adjacent frames from the video, processed with crop-align technique, to calculate optical flow features and emphasize specific facial areas by the Regions of Interest selection. These optical flow features are then fed to the three-stream convolution module for feature extraction. Subsequently, the extracted features are fed into the Swin-Transformer block and cross attention module for both intra and inter features interaction. Finally, MEs spotting on TCTFM seeks to predict a score indicating the probability that a frame falls within a certain expression range. The proposed spotting approach demonstrated exceptional performance in the MEGC 2021 spotting task, achieving the overall F1 score of 0.3547 and 0.3976 on the CAS(ME)2 dataset and the SAMM LV dataset, respectively. Experimental results show that MEs spotting based on TCTFM boost the representation capability on informative MEs details and outperforms existing MEs spotting models.