The main purpose of source-free cross-corpus speech emotion recognition (SER) is to use a pre-trained source model for emotion knowledge transfer. To address the challenge of inaccessible source data, we propose a simple yet effective method called local-global iterative adaptation network (LGIAN). The core idea is to enhance local aggregation and global discriminability via a meta-learning way. Specifically, we propose nearest neighbor clustering to capture intrinsic local structures in target data, ensuring local consistency. Supervised contrastive learning is utilized to maintain global emotion discriminability. We then alternate between nearest neighbour clustering and supervised contrastive learning as a meta-train and meta-test task, respectively, so that they can reinforce each other for better adaptation. We conduct extensive experiments on EmoDB, CASIA, eNTERFACE and EMOVO. Experimental results indicate the superior performance of LGIAN for source-free cross-corpus SER.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Local and Global Iterative Adaptation Based on Meta Learning for Source-Free Cross-Corpus Speech Emotion Recognition

  • Yan Zhao,
  • Jincen Wang,
  • Hailun Lian,
  • Sunan Li,
  • Jie Zhu,
  • Fan Liu

摘要

The main purpose of source-free cross-corpus speech emotion recognition (SER) is to use a pre-trained source model for emotion knowledge transfer. To address the challenge of inaccessible source data, we propose a simple yet effective method called local-global iterative adaptation network (LGIAN). The core idea is to enhance local aggregation and global discriminability via a meta-learning way. Specifically, we propose nearest neighbor clustering to capture intrinsic local structures in target data, ensuring local consistency. Supervised contrastive learning is utilized to maintain global emotion discriminability. We then alternate between nearest neighbour clustering and supervised contrastive learning as a meta-train and meta-test task, respectively, so that they can reinforce each other for better adaptation. We conduct extensive experiments on EmoDB, CASIA, eNTERFACE and EMOVO. Experimental results indicate the superior performance of LGIAN for source-free cross-corpus SER.