AraSER-Fusion: Multimodal Arabic Speech Emotion Recognition with Audio–Text Joint Embeddings
摘要
Arabic speech emotion recognition (SER) remains significantly underexplored, despite Arabic being spoken by over 400 million people and exhibiting substantial dialectal diversity that complicates affective modeling. Existing systems struggle with the linguistic, phonetic, and cultural variability across Arabic dialects, leading to limited robustness and weak generalization. This work introduces AraSER-Fusion, a multimodal framework designed to address these challenges by jointly modeling acoustic and textual signals within a dialect-aware architecture. AraSER-Fusion integrates transformer-based acoustic encoders to capture prosodic and spectral features, while AraBERT-derived embeddings extract contextual emotional cues from text. A hierarchical attention fusion module, supported by contrastive cross-modal alignment and adaptive gating, enables the model to effectively leverage complementary modality contributions. The system is optimized for the linguistic characteristics of Arabic and for real-world conditions. Comprehensive experiments across three diverse Arabic emotion datasets (AESC, KSA-Emotions, and MAED), representing Egyptian, Levantine, Gulf, Moroccan, and Algerian dialects, demonstrate that AraSER-Fusion achieves 87.3% accuracy and 86.1% F1-score, outperforming state-of-the-art methods by 8.5% and 9.2%, respectively. The model also maintains strong performance under high noise levels (80.4% accuracy at 5 dB SNR) and exhibits minimal performance variation across dialects (4.4% range). With 256.8 M parameters and 38.2 ms inference latency, the framework demonstrates suitability for practical deployment in human–computer interaction, mental health analysis, and culturally aligned virtual assistants.