Emotion recognition research in human–computer interactions has necessitated the development of automatic emotion identification systems. People’s ability to identify their emotions may help them better manage their emotions and engage with different situations in their lives. Numerous studies have explored into emotion categorization methods, although majority of them simply consider one, a small number, or independent physiological signs. Recognizing the importance of distinct signals and their integration will enable the development of further informative, economical, and objective techniques for detecting emotions, processing, and interpretations. In this paper, a novel FusionLSTM +  + is proposed that integrates multimodal inputs such as audio, text, and motion data, including facial expressions and hand movements by employing a hybrid neural network approach. It involves designing separate classification structures for each modality and fusing the outputs at the final layer, resulting in more reliable and accurate emotion detection. Experimental evaluation on the IEMOCAP dataset demonstrates the effectiveness of our neural network-based framework in capturing nuanced emotional expressions, contributing to the advancement of multimodal emotion recognition in human–computer interactions.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhanced Recognition of Human Emotion via Multimodal Inputs Using Fusion LSTM +  +  Model

  • S. Abinaya,
  • Annsley Mohan Joseph Raj,
  • Aishwarya Shaji

摘要

Emotion recognition research in human–computer interactions has necessitated the development of automatic emotion identification systems. People’s ability to identify their emotions may help them better manage their emotions and engage with different situations in their lives. Numerous studies have explored into emotion categorization methods, although majority of them simply consider one, a small number, or independent physiological signs. Recognizing the importance of distinct signals and their integration will enable the development of further informative, economical, and objective techniques for detecting emotions, processing, and interpretations. In this paper, a novel FusionLSTM +  + is proposed that integrates multimodal inputs such as audio, text, and motion data, including facial expressions and hand movements by employing a hybrid neural network approach. It involves designing separate classification structures for each modality and fusing the outputs at the final layer, resulting in more reliable and accurate emotion detection. Experimental evaluation on the IEMOCAP dataset demonstrates the effectiveness of our neural network-based framework in capturing nuanced emotional expressions, contributing to the advancement of multimodal emotion recognition in human–computer interactions.