<p>The identification of emotions has attracted significant analysis in various sectors, including mental health evaluation and human computer interaction. A limitation of emotion detection technology is its susceptibility to misread minor emotional signals, which might results in inaccurate results. Therefore, in this paper, an adaptive social spider optimized attentional gated recurrent unit (ASSO-AGRU) method is proposed that provides an exhaustive understanding of human emotions by integrating different modalities including voice, textual data and facial expressions. The multimodal dataset tokenization, Z-score normalization, and Gaussian filter (GF) for text, audio, and video are utilized through pre-processing of data. It is followed by extraction of features for Term Frequency - Inverse Document Frequency (TF-IDF), Mel-frequency Cepstral Coefficients (MFCCs), and histogram of oriented gradients (HOG) through text, audio, and video respectively. The ASSO-AGRU based integration strategy of multimodal inputs, which captures essential time-based relationships and correlations between different modalities, results in improved emotion recognition ability. The results show that the proposed framework attains F1-score (89%), precision (87.50%), recall (88.50%) and accuracy (89.70%) thus making it a reliable and logical method for multimodal emotion identification.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Emotion recognition with hybrid attentional multimodal fusion framework using cognitive augmentation

  • Shailesh Kulkarni,
  • S. S. Khot,
  • Yogesh Angal

摘要

The identification of emotions has attracted significant analysis in various sectors, including mental health evaluation and human computer interaction. A limitation of emotion detection technology is its susceptibility to misread minor emotional signals, which might results in inaccurate results. Therefore, in this paper, an adaptive social spider optimized attentional gated recurrent unit (ASSO-AGRU) method is proposed that provides an exhaustive understanding of human emotions by integrating different modalities including voice, textual data and facial expressions. The multimodal dataset tokenization, Z-score normalization, and Gaussian filter (GF) for text, audio, and video are utilized through pre-processing of data. It is followed by extraction of features for Term Frequency - Inverse Document Frequency (TF-IDF), Mel-frequency Cepstral Coefficients (MFCCs), and histogram of oriented gradients (HOG) through text, audio, and video respectively. The ASSO-AGRU based integration strategy of multimodal inputs, which captures essential time-based relationships and correlations between different modalities, results in improved emotion recognition ability. The results show that the proposed framework attains F1-score (89%), precision (87.50%), recall (88.50%) and accuracy (89.70%) thus making it a reliable and logical method for multimodal emotion identification.