<p>Facial expression recognition (FER) remains a challenging task due to subtle variations in facial details and unconstrained conditions such as changes in head posture, illumination, and occlusion. Current FER approaches primarily focus on capturing discriminative facial features in vision manner, often neglecting the rich semantic information available in textual modalities. Additionally, these methods typically rely on generic classification templates, which fail to capture instance-specific features, resulting in inadequate representation and fine-grained discrimination ability. To tackle the above issues, we propose a novel emotion-aware adaptation framework that integrates the pre-trained CLIP model for FER, leveraging both visual and textual modalities to enhance representation learning and capture fine-grained emotional details. Specifically, we introduce the Expression-aware adapter module to capture emotion-specific facial representations through task-specific fine-tuning while preserving the generalization capabilities of the CLIP model. Furthermore, the instance-enhanced expression classifier module is proposed to enhance textual descriptors with instance-specific visual embeddings using spherical linear interpolation, creating a more precise and discriminative classifier. Extensive experiments on three in-the-wild FER benchmarks demonstrate superiority of our proposed approach.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Emotion-aware adaptation of CLIP model for facial expression recognition

  • Jing Huan,
  • Mingxing Li,
  • Haoliang Zhou

摘要

Facial expression recognition (FER) remains a challenging task due to subtle variations in facial details and unconstrained conditions such as changes in head posture, illumination, and occlusion. Current FER approaches primarily focus on capturing discriminative facial features in vision manner, often neglecting the rich semantic information available in textual modalities. Additionally, these methods typically rely on generic classification templates, which fail to capture instance-specific features, resulting in inadequate representation and fine-grained discrimination ability. To tackle the above issues, we propose a novel emotion-aware adaptation framework that integrates the pre-trained CLIP model for FER, leveraging both visual and textual modalities to enhance representation learning and capture fine-grained emotional details. Specifically, we introduce the Expression-aware adapter module to capture emotion-specific facial representations through task-specific fine-tuning while preserving the generalization capabilities of the CLIP model. Furthermore, the instance-enhanced expression classifier module is proposed to enhance textual descriptors with instance-specific visual embeddings using spherical linear interpolation, creating a more precise and discriminative classifier. Extensive experiments on three in-the-wild FER benchmarks demonstrate superiority of our proposed approach.