Emotions play a critical role in human understanding and interpersonal communication. Deciphering emotions from text and audio sources presents significant challenges in Affective Computing and Human-Computer Interaction. The challenge intensifies for those with hearing or speech impairments, necessitating accessible technologies. Our study proposes a new approach, employing advanced AI to tackle text and audio data emotion recognition complexities. This approach directly caters to the unique needs of deaf and speech-impaired communities, striving to enhance the accuracy of emotion recognition and understanding. Our study innovates by integrating text and audio features via the CMU-Multimodal Opinion Sentiment and Emotion Intensity CMU-MOSI dataset, employing t-distributed Stochastic Neighbor Embedding (t-SNE) visualizations for comprehensive feature representation. This fusion approach aims to boost emotion detection accuracy by harnessing more information from each modality. Using CA-GNNs for classification, our method leverages the combined feature set for accurate emotion identification, thereby enhancing model effectiveness.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Emotion Recognition from Text and Audio Dataset Using Cumulative Attribute Graph Neural Networks (CA-GNN)

  • Hussein Farooq Tayeb Alsaadawi,
  • Resul Das

摘要

Emotions play a critical role in human understanding and interpersonal communication. Deciphering emotions from text and audio sources presents significant challenges in Affective Computing and Human-Computer Interaction. The challenge intensifies for those with hearing or speech impairments, necessitating accessible technologies. Our study proposes a new approach, employing advanced AI to tackle text and audio data emotion recognition complexities. This approach directly caters to the unique needs of deaf and speech-impaired communities, striving to enhance the accuracy of emotion recognition and understanding. Our study innovates by integrating text and audio features via the CMU-Multimodal Opinion Sentiment and Emotion Intensity CMU-MOSI dataset, employing t-distributed Stochastic Neighbor Embedding (t-SNE) visualizations for comprehensive feature representation. This fusion approach aims to boost emotion detection accuracy by harnessing more information from each modality. Using CA-GNNs for classification, our method leverages the combined feature set for accurate emotion identification, thereby enhancing model effectiveness.