<p>Recently, multimodal sentiment analysis (MSA) has gained significant traction due to its wide-range of applications in social media monitoring, healthcare, e-commerce, content creation, and business research. Unlike classical unimodal approaches, MSA integrates verbal and non-verbal characteristics such as text, speech, facial expressions, gestures, and physiological signals to enable a comprehensive understanding of human sentiments, particularly in human–computer interaction systems. Despite several existing reviews, many either focus on limited modalities or overlook the latest advancements, such as transformer-based and large language models (LLMs). This review presents a comprehensive and critical overview of MSA, with an integrated modalities, fusion techniques, classification, and current research in joint embedding and LLMs. It performs a comparative assessment of benchmark corpora, metrics, and model performance. The study also provides a critical analysis of current methods, highlighting their strengths, limitations, and future directions in the field. It also extends the discussion to practical applications and long-standing issues, which will frame the future research agenda in MSA. The Literature was carefully reviewed and selected from top academic databases using systematic search strategies. This survey aims to help researchers, students, and practitioners understand the history of MSA's development over the years and explore current research directions in the rapidly emerging field.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Systematic review of recent advances in multimodal sentiment analysis

  • Sumit Kumar Baberwal,
  • Nitin Arvind Shelke,
  • Khalid Anwar

摘要

Recently, multimodal sentiment analysis (MSA) has gained significant traction due to its wide-range of applications in social media monitoring, healthcare, e-commerce, content creation, and business research. Unlike classical unimodal approaches, MSA integrates verbal and non-verbal characteristics such as text, speech, facial expressions, gestures, and physiological signals to enable a comprehensive understanding of human sentiments, particularly in human–computer interaction systems. Despite several existing reviews, many either focus on limited modalities or overlook the latest advancements, such as transformer-based and large language models (LLMs). This review presents a comprehensive and critical overview of MSA, with an integrated modalities, fusion techniques, classification, and current research in joint embedding and LLMs. It performs a comparative assessment of benchmark corpora, metrics, and model performance. The study also provides a critical analysis of current methods, highlighting their strengths, limitations, and future directions in the field. It also extends the discussion to practical applications and long-standing issues, which will frame the future research agenda in MSA. The Literature was carefully reviewed and selected from top academic databases using systematic search strategies. This survey aims to help researchers, students, and practitioners understand the history of MSA's development over the years and explore current research directions in the rapidly emerging field.