In modern speech emotion recognition (SER), improving upon state-of-the-art systems requires increasing sophistication of extracting discriminative information from training data. The key aspects of our research include utilizing automatic speaker verification (ASV) and topic detection models to enhance emotion recognition in multimodal (audio and text) dialogues. We conduct a comparative analysis of modern SER models and experiment with various models for text (RoBERTa, TodKat, COSMIC) and audio (WavLM, Wav2Vec2) modalities. We also employ attention-based fusion of the modalities and a BiGRU-based classification approach. The results show that our proposed model, which combines RoBERTa, TodKat, Wav2Vec2, WavLM, and attention-based fusion, achieves an F1-score of 0.657, outperforming the state-of-the-art systems using the same selection of modalities. This performance is only surpassed by models that utilize additional modalities beyond audio and text. We demonstrate the effectiveness of our approach in improving SER by leveraging advanced techniques for extracting discriminative information from multimodal training data. The incorporation of ASV and topic detection, along with the novel fusion and classification methods, contributes to the enhanced performance of our proposed model compared to the existing state-of-the-art SER systems.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Utilizing Speaker Models and Topic Markers for Emotion Recognition in Dialogues

  • Olesia Makhnytkina,
  • Yuri Matveev,
  • Alexander Zubakov,
  • Anton Matveev

摘要

In modern speech emotion recognition (SER), improving upon state-of-the-art systems requires increasing sophistication of extracting discriminative information from training data. The key aspects of our research include utilizing automatic speaker verification (ASV) and topic detection models to enhance emotion recognition in multimodal (audio and text) dialogues. We conduct a comparative analysis of modern SER models and experiment with various models for text (RoBERTa, TodKat, COSMIC) and audio (WavLM, Wav2Vec2) modalities. We also employ attention-based fusion of the modalities and a BiGRU-based classification approach. The results show that our proposed model, which combines RoBERTa, TodKat, Wav2Vec2, WavLM, and attention-based fusion, achieves an F1-score of 0.657, outperforming the state-of-the-art systems using the same selection of modalities. This performance is only surpassed by models that utilize additional modalities beyond audio and text. We demonstrate the effectiveness of our approach in improving SER by leveraging advanced techniques for extracting discriminative information from multimodal training data. The incorporation of ASV and topic detection, along with the novel fusion and classification methods, contributes to the enhanced performance of our proposed model compared to the existing state-of-the-art SER systems.