A dual-interleaved feature differentiation network for mental health state recognition via facial expression analysis
摘要
With the growing concern over mental health, facial expression-based automatic monitoring has gained attention for its non-invasive and real-time advantages in emotion recognition and psychological assessment. However, existing methods often rely on single-level facial feature modeling, making them ineffective at handling the confusion caused by disguised or overlapping expressions, thus limiting their reliability in mental health recognition. To overcome these challenges, we propose the Dual-Interleaved Feature Differentiation Network (DIFNet) for differential perception and deep fusion of multi-granularity features. DIFNet employs a dual-branch architecture, separately modeling action unit (AU) features and global facial expressions, with a cross-attention mechanism for feature guidance and dynamic fusion. This design mitigates interference from disguised expressions. A Shunt Gating Unit (SGU) adaptively aggregates features, ensuring dominant expressions are effectively emphasized under complex conditions. We further introduce a Differential Graph Convolutional Network (DGCN), which enhances micro-expression differentiation through a differential amplification mechanism. To address feature distribution inconsistencies between branches, we adopt a KL divergence loss with channel-level semantic alignment, reinforcing collaboration and consistency in joint modeling. Experiments on AVEC2013 and AVEC2014 demonstrate that DIFNet outperforms existing methods across key metrics, showing enhanced robustness and mental state discrimination, particularly in the presence of disguised or composite emotions.