<p>With the rapid popularity of short video content, multimodal sentiment analysis (MSA) has attracted extensive attention. Most previous MSA studies have focused on manually transcribed benchmark datasets, which are both costly to generate and limited in availability. In real-world applications, MSA often relies on Automatic Speech Recognition (ASR) technology. However, due to the noise in the captions generated by ASR, traditional MSA models suffer a substantial decline in performance. To address this problem, this paper proposes a Multi-level Sentiment-aware Clustering Denoising Model (MSCDM), which effectively enhances the robustness by introducing sentiment distance constraints both intra- and inter-modality. Specifically, the model first compensates for the loss of sentiment semantic information in the text modality caused by ASR by leveraging samples with the same sentiment polarity to guide each other. Subsequently, the model refines cross-modal sentiment representations by dividing samples with multimodal information into positive and negative examples. We conduct extensive experiments on real-world datasets including MOSI-SpeechBrain, MOSI-IBM, and MOSI-iFlytek, and the results demonstrate the model’s effectiveness, outperforming the current state-of-the-art models on three datasets. The in-depth analysis confirms the performance of the multi-level clustering denoising strategy proposed for MSA.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-level sentiment-aware clustering for denoising in multimodal sentiment analysis with ASR errors

  • Zixu Hu,
  • Zhengtao Yu,
  • Junjun Guo

摘要

With the rapid popularity of short video content, multimodal sentiment analysis (MSA) has attracted extensive attention. Most previous MSA studies have focused on manually transcribed benchmark datasets, which are both costly to generate and limited in availability. In real-world applications, MSA often relies on Automatic Speech Recognition (ASR) technology. However, due to the noise in the captions generated by ASR, traditional MSA models suffer a substantial decline in performance. To address this problem, this paper proposes a Multi-level Sentiment-aware Clustering Denoising Model (MSCDM), which effectively enhances the robustness by introducing sentiment distance constraints both intra- and inter-modality. Specifically, the model first compensates for the loss of sentiment semantic information in the text modality caused by ASR by leveraging samples with the same sentiment polarity to guide each other. Subsequently, the model refines cross-modal sentiment representations by dividing samples with multimodal information into positive and negative examples. We conduct extensive experiments on real-world datasets including MOSI-SpeechBrain, MOSI-IBM, and MOSI-iFlytek, and the results demonstrate the model’s effectiveness, outperforming the current state-of-the-art models on three datasets. The in-depth analysis confirms the performance of the multi-level clustering denoising strategy proposed for MSA.