In the realm of Speech Emotion Recognition (SER), addressing the unique challenges presented by linguistically diverse and low-resource environments is critical yet often overlooked. This study advances SER for Urdu, a notably underrepresented language in computational linguistics, marking a significant stride toward linguistic inclusivity in intelligent systems. We introduce a pioneering combination of Mel-Frequency Cepstral Coefficients (MFCCs) with Convolutional Neural Networks (CNNs) and Gated Recurrent Unit (GRU) technologies to capture the emotional subtleties within Urdu speech. Our research employs the SEMOUR+ dataset for comprehensive analysis and incorporates cross-validation with an additional Urdu dataset using both GRU and Random Forest models to evaluate robustness. The results demonstrate a significant enhancement in SER accuracy, achieving up to 84.92%. Notably, the proficiency in identifying happiness as a test case highlights the model’s real-world applicability. This work not only furthers the development of SER frameworks for Urdu but also establishes a foundation for similar advancements in other low-resource languages, underscoring the crucial role of artificial intelligence in overcoming linguistic boundaries in emotion detection.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Advancing Speech Emotion Recognition for Urdu: Methodological Developments in Low-Resource Contexts

  • Muhammad Adeel,
  • Zhiyong Tao

摘要

In the realm of Speech Emotion Recognition (SER), addressing the unique challenges presented by linguistically diverse and low-resource environments is critical yet often overlooked. This study advances SER for Urdu, a notably underrepresented language in computational linguistics, marking a significant stride toward linguistic inclusivity in intelligent systems. We introduce a pioneering combination of Mel-Frequency Cepstral Coefficients (MFCCs) with Convolutional Neural Networks (CNNs) and Gated Recurrent Unit (GRU) technologies to capture the emotional subtleties within Urdu speech. Our research employs the SEMOUR+ dataset for comprehensive analysis and incorporates cross-validation with an additional Urdu dataset using both GRU and Random Forest models to evaluate robustness. The results demonstrate a significant enhancement in SER accuracy, achieving up to 84.92%. Notably, the proficiency in identifying happiness as a test case highlights the model’s real-world applicability. This work not only furthers the development of SER frameworks for Urdu but also establishes a foundation for similar advancements in other low-resource languages, underscoring the crucial role of artificial intelligence in overcoming linguistic boundaries in emotion detection.