<p>Arabic speech emotion recognition (SER) remains significantly underexplored, despite Arabic being spoken by over 400 million people and exhibiting substantial dialectal diversity that complicates affective modeling. Existing systems struggle with the linguistic, phonetic, and cultural variability across Arabic dialects, leading to limited robustness and weak generalization. This work introduces AraSER-Fusion, a multimodal framework designed to address these challenges by jointly modeling acoustic and textual signals within a dialect-aware architecture. AraSER-Fusion integrates transformer-based acoustic encoders to capture prosodic and spectral features, while AraBERT-derived embeddings extract contextual emotional cues from text. A hierarchical attention fusion module, supported by contrastive cross-modal alignment and adaptive gating, enables the model to effectively leverage complementary modality contributions. The system is optimized for the linguistic characteristics of Arabic and for real-world conditions. Comprehensive experiments across three diverse Arabic emotion datasets (AESC, KSA-Emotions, and MAED), representing Egyptian, Levantine, Gulf, Moroccan, and Algerian dialects, demonstrate that AraSER-Fusion achieves 87.3% accuracy and 86.1% F1-score, outperforming state-of-the-art methods by 8.5% and 9.2%, respectively. The model also maintains strong performance under high noise levels (80.4% accuracy at 5&#xa0;dB SNR) and exhibits minimal performance variation across dialects (4.4% range). With 256.8&#xa0;M parameters and 38.2&#xa0;ms inference latency, the framework demonstrates suitability for practical deployment in human–computer interaction, mental health analysis, and culturally aligned virtual assistants.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

AraSER-Fusion: Multimodal Arabic Speech Emotion Recognition with Audio–Text Joint Embeddings

  • Reham Al-Dayil

摘要

Arabic speech emotion recognition (SER) remains significantly underexplored, despite Arabic being spoken by over 400 million people and exhibiting substantial dialectal diversity that complicates affective modeling. Existing systems struggle with the linguistic, phonetic, and cultural variability across Arabic dialects, leading to limited robustness and weak generalization. This work introduces AraSER-Fusion, a multimodal framework designed to address these challenges by jointly modeling acoustic and textual signals within a dialect-aware architecture. AraSER-Fusion integrates transformer-based acoustic encoders to capture prosodic and spectral features, while AraBERT-derived embeddings extract contextual emotional cues from text. A hierarchical attention fusion module, supported by contrastive cross-modal alignment and adaptive gating, enables the model to effectively leverage complementary modality contributions. The system is optimized for the linguistic characteristics of Arabic and for real-world conditions. Comprehensive experiments across three diverse Arabic emotion datasets (AESC, KSA-Emotions, and MAED), representing Egyptian, Levantine, Gulf, Moroccan, and Algerian dialects, demonstrate that AraSER-Fusion achieves 87.3% accuracy and 86.1% F1-score, outperforming state-of-the-art methods by 8.5% and 9.2%, respectively. The model also maintains strong performance under high noise levels (80.4% accuracy at 5 dB SNR) and exhibits minimal performance variation across dialects (4.4% range). With 256.8 M parameters and 38.2 ms inference latency, the framework demonstrates suitability for practical deployment in human–computer interaction, mental health analysis, and culturally aligned virtual assistants.