<p>Facial Expression Recognition in natural scenes constitutes a critical research direction in affective computing. Prevailing approaches predominantly rely on RGB spatial domain modeling, neglecting the inherent texture prior information in the frequency domain, which consequently limits their adaptability to challenging scenarios. Furthermore, conventional methods typically depend on single-domain spatial features, failing to effectively exploit the complementary characteristics between spatial and frequency domains, thereby limiting the model’s capacity for subtle expression representation. To address these limitations, this paper proposes FPNet, a dual-domain collaborative framework for FER. Specifically, we design three core components: (1) DyLKBlock constructs dynamic spatial feature extraction through cascaded large-kernel convolutions, achieving an equivalent 23<InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mo>×</mo> </math></EquationSource> </InlineEquation>23 receptive fields while balancing global modeling capability with linear computational complexity; (2) FPBlock utilizes <b>D</b>iscrete <b>W</b>avelet <b>T</b>ransform <b>(DWT)</b> to achieve lossless downsampling, preserving multi-scale texture details; (3) CAFusion facilitates bidirectional cross-domain feature interaction through an attention mechanism, ensuring the preservation of discriminative frequency domain characteristics in cross-domain data. Extensive experiments on three benchmark datasets demonstrate that FPNet significantly outperforms SOTA methods. Visualization analyses further validate that the frequency domain priors in FPNet effectively capture subtle expressions in complex scenarios, with precise focus on key facial regions.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Frequency-prior enhanced network for facial expression recognition via dynamic large kernels and dual-domain learning

  • Chuanyu Cai,
  • Ke Chen

摘要

Facial Expression Recognition in natural scenes constitutes a critical research direction in affective computing. Prevailing approaches predominantly rely on RGB spatial domain modeling, neglecting the inherent texture prior information in the frequency domain, which consequently limits their adaptability to challenging scenarios. Furthermore, conventional methods typically depend on single-domain spatial features, failing to effectively exploit the complementary characteristics between spatial and frequency domains, thereby limiting the model’s capacity for subtle expression representation. To address these limitations, this paper proposes FPNet, a dual-domain collaborative framework for FER. Specifically, we design three core components: (1) DyLKBlock constructs dynamic spatial feature extraction through cascaded large-kernel convolutions, achieving an equivalent 23 \(\times \) × 23 receptive fields while balancing global modeling capability with linear computational complexity; (2) FPBlock utilizes Discrete Wavelet Transform (DWT) to achieve lossless downsampling, preserving multi-scale texture details; (3) CAFusion facilitates bidirectional cross-domain feature interaction through an attention mechanism, ensuring the preservation of discriminative frequency domain characteristics in cross-domain data. Extensive experiments on three benchmark datasets demonstrate that FPNet significantly outperforms SOTA methods. Visualization analyses further validate that the frequency domain priors in FPNet effectively capture subtle expressions in complex scenarios, with precise focus on key facial regions.