<p>In real-world mobile environments, the performance of Automatic Speech Recognition (ASR) systems is significantly compromised by environmental noise and distortions introduced by speech codecs. This study investigates two key components for enhancing ASR robustness: speech denoising and acoustic feature extraction. We explore two distinct feature extraction techniques: the Gabor filter bank, which improves time–frequency resolution, and Multi-Resolution Cochleagram (MRCG) features derived from the Mel spectrum, designed to enhance perceptual representation. For speech denoising, two advanced approaches are evaluated: Non-negative Matrix Factorization (NMF) and Computational Auditory Scene Analysis (CASA). The latter utilizes Ideal Binary Masking (IBM), estimated through Deep Neural Networks (DNN), to effectively isolate target speech from background noise. The ASR system is based on a Convolutional Neural Network (CNN) enhanced with an attention mechanism, enabling more accurate and robust speech recognition in noisy mobile environments Experimentation is conducted on the ARADIGIT database in quiet and noisy settings with various Signal-to-Noise Ratios (SNRs), including babble noise ranging from − 10&#xa0;dB to 15&#xa0;dB. Results show that combining IBM-DNN separation with The hybrid (MRCG + MFCC + ∆ + ∆∆) front-end demonstrates the best generalization capability across noise levels achieving the highest recognition accuracy, reaching up to 98%, supporting its suitability for robust automatic speech recognition tasks. In contrast, conventional feature combinations (named as Global features for: SGBFB + MRCG + MFCC + ∆ + ∆∆) with NMF yield lower maximum accuracies of around 78%. These findings highlight the effectiveness of integrating advanced denoising and feature extraction techniques, particularly IBM-DNN with (MRCG + MFCC + ∆ + ∆∆), in significantly improving ASR performance under adverse acoustic conditions. The study confirms the critical role of such techniques in enhancing ASR resilience and accuracy in real-world noisy environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Robustness of Arabic Speech Recognition in Noisy Environments Using Advanced Feature Extraction and Denoising Techniques Based on Deep Learning Models

  • Lallouani Bouchakour,
  • Khaled Lounnas,
  • Mohamed Debyeche

摘要

In real-world mobile environments, the performance of Automatic Speech Recognition (ASR) systems is significantly compromised by environmental noise and distortions introduced by speech codecs. This study investigates two key components for enhancing ASR robustness: speech denoising and acoustic feature extraction. We explore two distinct feature extraction techniques: the Gabor filter bank, which improves time–frequency resolution, and Multi-Resolution Cochleagram (MRCG) features derived from the Mel spectrum, designed to enhance perceptual representation. For speech denoising, two advanced approaches are evaluated: Non-negative Matrix Factorization (NMF) and Computational Auditory Scene Analysis (CASA). The latter utilizes Ideal Binary Masking (IBM), estimated through Deep Neural Networks (DNN), to effectively isolate target speech from background noise. The ASR system is based on a Convolutional Neural Network (CNN) enhanced with an attention mechanism, enabling more accurate and robust speech recognition in noisy mobile environments Experimentation is conducted on the ARADIGIT database in quiet and noisy settings with various Signal-to-Noise Ratios (SNRs), including babble noise ranging from − 10 dB to 15 dB. Results show that combining IBM-DNN separation with The hybrid (MRCG + MFCC + ∆ + ∆∆) front-end demonstrates the best generalization capability across noise levels achieving the highest recognition accuracy, reaching up to 98%, supporting its suitability for robust automatic speech recognition tasks. In contrast, conventional feature combinations (named as Global features for: SGBFB + MRCG + MFCC + ∆ + ∆∆) with NMF yield lower maximum accuracies of around 78%. These findings highlight the effectiveness of integrating advanced denoising and feature extraction techniques, particularly IBM-DNN with (MRCG + MFCC + ∆ + ∆∆), in significantly improving ASR performance under adverse acoustic conditions. The study confirms the critical role of such techniques in enhancing ASR resilience and accuracy in real-world noisy environments.