Advancing Speech Emotion Recognition: A Comparative Study and Enhanced Performance in English and Algerian Dialect Using Uni and Multimodal Transformer
摘要
In this study, we conduct a comprehensive comparative analysis of various deep learning techniques for the Speech Emotion Recognition (SER) task, ranging from LSTM to CNN-BiLSTM-Multihead-Attention. Our evaluation utilizes the CREMA-D dataset for English and a crowdsourced dataset for the Algerian dialect. We propose two novel transformer-based frameworks, including an unimodal Transformer model and a multimodal Transformer model integrating DziriBERT. The Transformer-based model achieved an accuracy of 95% on the CREMA-D dataset, while the multimodal Transformer model with DziriBERT attained an impressive 97.42% accuracy on the Algerian dialect dataset. The proposed frameworks, unimodal Transformer and multimodal Transformer+DziriBERT significantly outperformed previous studies that used the same employed datasets in English language and Algerian dialect, respectively.