<p>In the micro-expression recognition task, the short duration of micro-expressions and the synergistic changes involving multiple muscle groups can affect the comprehensiveness of facial feature extraction. The dominant micro-expression feature extraction methods only rely on the motion information in the optical flow feature map and may ignore the temporal features of micro-expressions. It is also easy to ignore the overall correlation of facial expressions. Consequently, we propose a Dual Stream Network with embedding Temporal Convolution (DSNTC), which integrates temporal motion information with local and global information to achieve comprehensive feature extraction. In this model, we use temporal convolution operations to enhance the RepViT model. Considering the extraction of local and global information, we combine the MobileViT model with the RepViT model. We fuse the features through Cross-Attention Fusion and finally construct the DSNTC. Experimental results show that the accuracy of the model on CASME II, SAMM, and SMIC three public datasets reaches 89.21%, 78.76%, and 75.31%, respectively, indicating that the model has superior performance in recognition accuracy, complexity, and model parameters. The codes and models are available at: <a href="https://github.com/HQwilbur/DSNTC-Code">https://github.com/HQwilbur/DSNTC-Code</a></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The dual stream network with embedding temporal convolution for micro-expression recognition

  • Haiquan Wang,
  • Kunxia Wang,
  • Wancheng Yu

摘要

In the micro-expression recognition task, the short duration of micro-expressions and the synergistic changes involving multiple muscle groups can affect the comprehensiveness of facial feature extraction. The dominant micro-expression feature extraction methods only rely on the motion information in the optical flow feature map and may ignore the temporal features of micro-expressions. It is also easy to ignore the overall correlation of facial expressions. Consequently, we propose a Dual Stream Network with embedding Temporal Convolution (DSNTC), which integrates temporal motion information with local and global information to achieve comprehensive feature extraction. In this model, we use temporal convolution operations to enhance the RepViT model. Considering the extraction of local and global information, we combine the MobileViT model with the RepViT model. We fuse the features through Cross-Attention Fusion and finally construct the DSNTC. Experimental results show that the accuracy of the model on CASME II, SAMM, and SMIC three public datasets reaches 89.21%, 78.76%, and 75.31%, respectively, indicating that the model has superior performance in recognition accuracy, complexity, and model parameters. The codes and models are available at: https://github.com/HQwilbur/DSNTC-Code