<p>Emotion recognition has emerged as a pivotal area of research within artificial intelligence (AI) and human–computer interaction, enabling computational systems to perceive, interpret, and respond to human affective states. Despite substantial advancements, challenges related to cross-domain variability and multimodal inconsistency continue to hinder the development of robust and generalizable emotion classification models. This study proposes an explainable cross-domain emotion recognition framework that integrates advanced preprocessing, multimodal fusion, and adaptive optimization techniques. The proposed model leverages Bidirectional Encoder Representations from Transformers (BERT) for textual data, ResNet-50 for visual features, and Mel-Frequency Cepstral Coefficients (MFCC) for acoustic representations. A Long Short-Term Memory (LSTM) network serves as the core classifier, with its parameters optimized via the Reptile Search Algorithm (RSA), a bio-inspired optimization technique that enhances exploration, convergence, and training stability. This LSTM-RSA configuration effectively captures temporal dependencies and cross-modal correlations, thereby improving emotion recognition accuracy across diverse scenarios. Furthermore, Shapley Additive Explanations (SHAP) are incorporated to provide interpretability by quantifying the contribution of each feature, thereby strengthening user trust and system transparency. Experimental validation on benchmark multimodal datasets, IEMOCAP and SAVEE, demonstrates superior performance, achieving accuracies of 98.12% and 97.86%, respectively.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Explainable cross-domain emotion recognition using non-linear optimization and multimodal feature fusion based deep learning model

  • Rami Baazeem

摘要

Emotion recognition has emerged as a pivotal area of research within artificial intelligence (AI) and human–computer interaction, enabling computational systems to perceive, interpret, and respond to human affective states. Despite substantial advancements, challenges related to cross-domain variability and multimodal inconsistency continue to hinder the development of robust and generalizable emotion classification models. This study proposes an explainable cross-domain emotion recognition framework that integrates advanced preprocessing, multimodal fusion, and adaptive optimization techniques. The proposed model leverages Bidirectional Encoder Representations from Transformers (BERT) for textual data, ResNet-50 for visual features, and Mel-Frequency Cepstral Coefficients (MFCC) for acoustic representations. A Long Short-Term Memory (LSTM) network serves as the core classifier, with its parameters optimized via the Reptile Search Algorithm (RSA), a bio-inspired optimization technique that enhances exploration, convergence, and training stability. This LSTM-RSA configuration effectively captures temporal dependencies and cross-modal correlations, thereby improving emotion recognition accuracy across diverse scenarios. Furthermore, Shapley Additive Explanations (SHAP) are incorporated to provide interpretability by quantifying the contribution of each feature, thereby strengthening user trust and system transparency. Experimental validation on benchmark multimodal datasets, IEMOCAP and SAVEE, demonstrates superior performance, achieving accuracies of 98.12% and 97.86%, respectively.