In human-computer interaction (HCI), speech emotion recognition (SER) is a key technology that enables machines to detect emotions in voices for more effective communication. Recently, pre-trained speech representations have demonstrated significant potential for SER. However, due to the inherent variability in individual emotional expressions, models often focus more on identity information than emotional information within these representations. This can hinder the ability to fully leverage the capabilities of pre-trained representations for emotion recognition. To address this challenge, we propose a phoneme-aware multi-task learning framework tailored for SER. Specifically, the framework incorporates phoneme recognition as an auxiliary task and leverages the speaker-invariance of phoneme articulation to reduce the influence of identity information. Additionally, a dynamic task prioritization strategy has been customized to tackle the imbalance during the multi-task training process, leveraging the synergistic effects between tasks. To further enhance the utilization of pre-trained representations, we introduce a Squeeze-and-Excitation module to extract emotion-related information. Experimental results on the IEMOCAP and MELD datasets demonstrate that our approach achieves state-of-the-art performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Phoneme-Aware Multi-task Learning Framework with Dynamic Prioritization for Speech Emotion Recognition

  • Kang Yin,
  • Sirui Zhao,
  • Xinlong Mao,
  • Shifeng Liu,
  • Yiming Zhang,
  • Tong Xu,
  • Enhong Chen

摘要

In human-computer interaction (HCI), speech emotion recognition (SER) is a key technology that enables machines to detect emotions in voices for more effective communication. Recently, pre-trained speech representations have demonstrated significant potential for SER. However, due to the inherent variability in individual emotional expressions, models often focus more on identity information than emotional information within these representations. This can hinder the ability to fully leverage the capabilities of pre-trained representations for emotion recognition. To address this challenge, we propose a phoneme-aware multi-task learning framework tailored for SER. Specifically, the framework incorporates phoneme recognition as an auxiliary task and leverages the speaker-invariance of phoneme articulation to reduce the influence of identity information. Additionally, a dynamic task prioritization strategy has been customized to tackle the imbalance during the multi-task training process, leveraging the synergistic effects between tasks. To further enhance the utilization of pre-trained representations, we introduce a Squeeze-and-Excitation module to extract emotion-related information. Experimental results on the IEMOCAP and MELD datasets demonstrate that our approach achieves state-of-the-art performance.