This study presents a novel approach to Arabic speech command recognition using a hybrid Convolutional Neural Network-Long Short-Term Memory (CNN-LSTM) architecture. The research utilizes the Arabic Speech Commands dataset, comprising 12,000 audio samples of 40 unique keywords recorded by 30 participants. We implement mel frequency cepstral coefficients (MFCC) for feature extraction and develop a sophisticated model architecture combining 1D convolutional layers with causal padding and LSTM networks for temporal sequence modeling. The model incorporates multiple convolutional layers with increasing dilation rates (2, 4, 8, 11) and shortcut connections, followed by dual LSTM layers for capturing long-term dependencies. Experimental results demonstrate exceptional performance, achieving 99% overall accuracy on the validation dataset, with weighted average precision of 98.93, recall of 98.85, and F1-score of 98.84. The confusion matrix analysis reveals minimal misclassifications (1.15% error rate), primarily between acoustically similar commands. Learning curve analysis confirms robust generalization capabilities without overfitting, suggesting the model’s suitability for practical applications in Arabic speech command recognition systems. This research contributes to advancing Arabic speech recognition technology by demonstrating the effectiveness of a hybrid deep learning architecture in handling the unique challenges of Arabic speech processing.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Advanced Arabic Speech Recognition Through Dilated Convolutional LSTM Networks

  • Ezzaldeen Mahyoub Naji,
  • Ajit A. Maslekar,
  • Zeyad A. T. Ahmed,
  • Belal Al-sellami

摘要

This study presents a novel approach to Arabic speech command recognition using a hybrid Convolutional Neural Network-Long Short-Term Memory (CNN-LSTM) architecture. The research utilizes the Arabic Speech Commands dataset, comprising 12,000 audio samples of 40 unique keywords recorded by 30 participants. We implement mel frequency cepstral coefficients (MFCC) for feature extraction and develop a sophisticated model architecture combining 1D convolutional layers with causal padding and LSTM networks for temporal sequence modeling. The model incorporates multiple convolutional layers with increasing dilation rates (2, 4, 8, 11) and shortcut connections, followed by dual LSTM layers for capturing long-term dependencies. Experimental results demonstrate exceptional performance, achieving 99% overall accuracy on the validation dataset, with weighted average precision of 98.93, recall of 98.85, and F1-score of 98.84. The confusion matrix analysis reveals minimal misclassifications (1.15% error rate), primarily between acoustically similar commands. Learning curve analysis confirms robust generalization capabilities without overfitting, suggesting the model’s suitability for practical applications in Arabic speech command recognition systems. This research contributes to advancing Arabic speech recognition technology by demonstrating the effectiveness of a hybrid deep learning architecture in handling the unique challenges of Arabic speech processing.