Imagine someone talking on a busy street and their words suppressed by chatter and traffic sound. Our research sets out to tackle such noisy audio recordings and how CNNs can rise to this challenge, restoring clarity to communication. This paper presents an innovative pipeline for the enhancement of speech, that focuses on improving the quality of speech signals which are corrupted by background noise. Our method combines the Denoising Autoencoder (DAE) and Convolutional Neural Network (CNN) which is integrated with classical signal processing techniques. When trained on a sufficient amount of noisy speech frames, the DAE outperforms traditional ideal ratio mask (IRM) approaches when it comes to estimating in clean speech spectrum, resulting denoised speech is analyzed by using the STFTShor (t-Time Fourier Transform) and then reconstructed using the overlap-add technique, providing a refined output. This entire pipeline can adapt to diverse acoustic scenarios, thus presenting an all-purpose solution for real-world applications starting from communication systems to voice- activated technologies. Our research demonstrates exceptional performance and a significant improvement in speech enhancement. The solution of integrating DAE-CNN answers existing difficulties in noise reduction and also contributes significantly to Machine learning advancements related to audio processing.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Audio Denoising: Speech Enhancement Using Convolutional Neural Networks

  • T. Aruna,
  • Shreya Vaidya,
  • Rakshita Raikar,
  • Satish Chikkamath

摘要

Imagine someone talking on a busy street and their words suppressed by chatter and traffic sound. Our research sets out to tackle such noisy audio recordings and how CNNs can rise to this challenge, restoring clarity to communication. This paper presents an innovative pipeline for the enhancement of speech, that focuses on improving the quality of speech signals which are corrupted by background noise. Our method combines the Denoising Autoencoder (DAE) and Convolutional Neural Network (CNN) which is integrated with classical signal processing techniques. When trained on a sufficient amount of noisy speech frames, the DAE outperforms traditional ideal ratio mask (IRM) approaches when it comes to estimating in clean speech spectrum, resulting denoised speech is analyzed by using the STFTShor (t-Time Fourier Transform) and then reconstructed using the overlap-add technique, providing a refined output. This entire pipeline can adapt to diverse acoustic scenarios, thus presenting an all-purpose solution for real-world applications starting from communication systems to voice- activated technologies. Our research demonstrates exceptional performance and a significant improvement in speech enhancement. The solution of integrating DAE-CNN answers existing difficulties in noise reduction and also contributes significantly to Machine learning advancements related to audio processing.