<p>Speech Emotion Recognition (SER) aims to identify human emotions from vocal expressions, contributing significantly to affective computing applications. Despite progress in SER for widely spoken languages, Urdu remains underrepresented due to the lack of dedicated datasets and language-specific modeling. This study proposes a deep learning framework specifically designed for Urdu SER, utilizing robust acoustic features including Mel Frequency Cepstral Coefficients (MFCC) with delta and double-delta derivatives, as well as Mel-spectrograms, to capture both spectral and temporal aspects of speech. We test and analyze three different model architectures: a standalone Convolutional Neural Network (CNN) with attention mechanism, a CNN-LSTM hybrid, and a CNN-LSTM hybrid with an attention mechanism. The attention-based CNN-LSTM model had the highest accuracy 95%, efficiently capturing subtle emotional cues in Urdu speech. These results highlight the potential of attention mechanisms and hybrid architectures in advancing emotion recognition for under-resourced languages like Urdu.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Emotion detection in Urdu speech: a deep hybrid learning approach

  • Muhammad Owais,
  • Khadija Sarwar,
  • Khurram Khan Jadoon,
  • Junaid Yousaf

摘要

Speech Emotion Recognition (SER) aims to identify human emotions from vocal expressions, contributing significantly to affective computing applications. Despite progress in SER for widely spoken languages, Urdu remains underrepresented due to the lack of dedicated datasets and language-specific modeling. This study proposes a deep learning framework specifically designed for Urdu SER, utilizing robust acoustic features including Mel Frequency Cepstral Coefficients (MFCC) with delta and double-delta derivatives, as well as Mel-spectrograms, to capture both spectral and temporal aspects of speech. We test and analyze three different model architectures: a standalone Convolutional Neural Network (CNN) with attention mechanism, a CNN-LSTM hybrid, and a CNN-LSTM hybrid with an attention mechanism. The attention-based CNN-LSTM model had the highest accuracy 95%, efficiently capturing subtle emotional cues in Urdu speech. These results highlight the potential of attention mechanisms and hybrid architectures in advancing emotion recognition for under-resourced languages like Urdu.