<p>Micro-expressions(MEs) have emerged as a viable strategy for affective estimation due to their high reliability in emotion detection. In recent years, deep learning methods have been successfully applied to the field of micro-expression recognition. However, extracting and learning features from MEs presents challenges due to their brief duration and subtle intensity. To address these challenges, we propose the dual-stream fusion network (DSFNet). Specifically, we design shallow tokens-to-token vision transformers (T2T-ViT) to effectively capture comprehensive spatial position information. We also fine-tuned the number of ViT encoders and heads to enhance overall model performance. Additionally, the proposed multiscale convolution block (MCB) and attention mechanism modules (AMM) facilitate the effective extraction of detailed and valuable multiscale features from MEs. By employing various sizes of convolutional kernels and attention mechanisms, our approach captures higher-level image information, thereby improving MER accuracy. Finally, we integrate the information obtained from both branches. Performance evaluations on three mainstream ME datasets-SMIC, CASME II, and SAMM-demonstrate that the proposed framework significantly outperforms other advanced methods in micro-expression classification.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Facial micro-expression recognition based on dual-stream fusion network

  • Jiacheng Sun,
  • Changhong Chen

摘要

Micro-expressions(MEs) have emerged as a viable strategy for affective estimation due to their high reliability in emotion detection. In recent years, deep learning methods have been successfully applied to the field of micro-expression recognition. However, extracting and learning features from MEs presents challenges due to their brief duration and subtle intensity. To address these challenges, we propose the dual-stream fusion network (DSFNet). Specifically, we design shallow tokens-to-token vision transformers (T2T-ViT) to effectively capture comprehensive spatial position information. We also fine-tuned the number of ViT encoders and heads to enhance overall model performance. Additionally, the proposed multiscale convolution block (MCB) and attention mechanism modules (AMM) facilitate the effective extraction of detailed and valuable multiscale features from MEs. By employing various sizes of convolutional kernels and attention mechanisms, our approach captures higher-level image information, thereby improving MER accuracy. Finally, we integrate the information obtained from both branches. Performance evaluations on three mainstream ME datasets-SMIC, CASME II, and SAMM-demonstrate that the proposed framework significantly outperforms other advanced methods in micro-expression classification.