<p>Multimodal sentiment analysis (MSA), which is aiming to comprehend the sentiments expressed through multimedia content, necessitates robust models able to effectively integrate data extracted from diverse modalities. The explosion of multimedia content on social media platforms in recent years has resulted in growing demand for advanced sentiment analysis methods that can efficiently handle various data modalities. This paper presents a novel deep learning-driven model for MSA named RoLiVit, which combines textual, auditory, and visual cues to capture the complex emotional expressions found in multimedia content. The proposed framework RoLiVit uses an ordered structure in which Roberta, Librosa, and Vision Transformer are used to first extract features from each modality. Long Short-Term Memory (LSTM), a type of recurrent neural network, then combines these modality-specific features to efficiently capture the complementary information found in each modality. In order to bolster these developments, assessment results on standard multimodal sentiment analysis datasets show that the suggested model performs better than the most advanced methods, attaining higher accuracy and adaptability on a variety of multimedia content types. The model was tested using the CMU-MOSI dataset, where it beat earlier state-of-the-art models by 1.70%, achieving 91% accuracy and an F1-score of 90.99%. These findings show how well the model performs on MSA tasks by capturing minute emotional details in several modalities. This approach employs the strengths of each modality (textual, audio, visual) to present a broader comprehension of sentiment in various forms of media, eventually leading to improved evaluation of sentiment performance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

RoLiVit: Feature Fusion Approach for Multimodal Sentiment Analysis Using Deep Learning

  • Namrata Shroff,
  • Shreya Patel,
  • Hemani Shah

摘要

Multimodal sentiment analysis (MSA), which is aiming to comprehend the sentiments expressed through multimedia content, necessitates robust models able to effectively integrate data extracted from diverse modalities. The explosion of multimedia content on social media platforms in recent years has resulted in growing demand for advanced sentiment analysis methods that can efficiently handle various data modalities. This paper presents a novel deep learning-driven model for MSA named RoLiVit, which combines textual, auditory, and visual cues to capture the complex emotional expressions found in multimedia content. The proposed framework RoLiVit uses an ordered structure in which Roberta, Librosa, and Vision Transformer are used to first extract features from each modality. Long Short-Term Memory (LSTM), a type of recurrent neural network, then combines these modality-specific features to efficiently capture the complementary information found in each modality. In order to bolster these developments, assessment results on standard multimodal sentiment analysis datasets show that the suggested model performs better than the most advanced methods, attaining higher accuracy and adaptability on a variety of multimedia content types. The model was tested using the CMU-MOSI dataset, where it beat earlier state-of-the-art models by 1.70%, achieving 91% accuracy and an F1-score of 90.99%. These findings show how well the model performs on MSA tasks by capturing minute emotional details in several modalities. This approach employs the strengths of each modality (textual, audio, visual) to present a broader comprehension of sentiment in various forms of media, eventually leading to improved evaluation of sentiment performance.