Selection of Preprocessing Parameters for Wave-U-Net-Based Speech-Denoising
摘要
Reducing unwanted noise to improve speech quality and intelligibility in a corrupted recording is challenging, especially when the noise is non-stationary or high level. Deep learning models, which have gained popularity in recent years, perform better under these challenging, noisy conditions than conventional approaches involving filtering, passive and active noise cancellation, or similar methods. However, deep learning-based studies usually focus on the network architecture and its optimization. This research investigated the impact of different signal parameters, namely signal sampling rate and analysis frame length, on the performance of a denoising model. A waveform-based network, Wave-U-Net, was used. Experimental results using sampling rates of 8, 11.025, 16, and 48 kHz and frame lengths between 10 and 50 ms showed that lower sampling rates produce better results. The 8 kHz sampling rate resulted in the highest PESQ score of 3.24, and the 11.025 kHz sampling rate resulted in the highest SNR of 15.77 dB. Both the 8 kHz and 11.025 kHz sampling rates resulted in the lowest MSE of 1.5 × \({10}^{-4}\) . The effective frame length for the lower sampling rates was 30–40 ms, while the optimum frame length could not be determined for the higher sampling rates.