Recent advancements in deep learning have significantly enhanced facial video-based depression recognition. However, existing models encounter critical limitations that hinder their performance. They struggle with spatially localized facial feature extraction, relying on either a single convolution or a convolution operation with a simplistic attention mechanism, leading to inadequate recognition of depression-relevant patterns. Besides, increasing model depth leads to the extraction of ambiguous, abstract features while losing crucial dynamic facial details. To address these issues, we propose LMS-VDR by combining the Multi-Scale Mixed Attention Module (MSMAM) with landmark-based prior knowledge integration. More specifically, MSMAM synergistically merges channel and spatial attention, forming a mixed attention vector block through vector products. It introduces a dense connection mechanism, directly connecting features of each dimension to the final output, thereby enabling multi-scale diversity feature extraction. The integration of landmarks, initially linearly transformed and later combined with temporal feature sequences, enhances dynamic temporal feature extraction using our proposed Cross Multi-head Self-Attention (CMHSA) block based on self-attention. Experiments on AVEC 2013 and AVEC 2014 datasets validate our method’s efficacy, achieving MAE/RMSE of 6.04/7.68 and 5.98/7.59, respectively. Our proposed method offers a promising direction for clinical depression assessment, demonstrating the potential for significant contributions in this critical domain.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LMS-VDR: Integrating Landmarks into Multi-scale Hybrid Net for Video-Based Depression Recognition

  • Mengyuan Yang,
  • Yuanyuan Shang,
  • Jingyi Liu,
  • Zhuhong Shao,
  • Tie Liu,
  • Hui Ding,
  • Hailiang Li

摘要

Recent advancements in deep learning have significantly enhanced facial video-based depression recognition. However, existing models encounter critical limitations that hinder their performance. They struggle with spatially localized facial feature extraction, relying on either a single convolution or a convolution operation with a simplistic attention mechanism, leading to inadequate recognition of depression-relevant patterns. Besides, increasing model depth leads to the extraction of ambiguous, abstract features while losing crucial dynamic facial details. To address these issues, we propose LMS-VDR by combining the Multi-Scale Mixed Attention Module (MSMAM) with landmark-based prior knowledge integration. More specifically, MSMAM synergistically merges channel and spatial attention, forming a mixed attention vector block through vector products. It introduces a dense connection mechanism, directly connecting features of each dimension to the final output, thereby enabling multi-scale diversity feature extraction. The integration of landmarks, initially linearly transformed and later combined with temporal feature sequences, enhances dynamic temporal feature extraction using our proposed Cross Multi-head Self-Attention (CMHSA) block based on self-attention. Experiments on AVEC 2013 and AVEC 2014 datasets validate our method’s efficacy, achieving MAE/RMSE of 6.04/7.68 and 5.98/7.59, respectively. Our proposed method offers a promising direction for clinical depression assessment, demonstrating the potential for significant contributions in this critical domain.