Public telephone systems are bandwidth-limited to 0.3 \(\sim \) 3.4 kHz due to telephone communication regulations, which significantly reduces voice quality and intelligibility. However, on the receiver side, it is possible to improve the clarity and intelligibility of the output speech quality by efficiently estimating and adding the high-frequency components lost from bandwidth limitation using only the narrow-band components. In this study, we use machine learning to estimate the high-frequency components in the range of 3.4 \(\sim \) 7 kHz that are normally lost due to bandwidth limitation, by utilizing the correlation between each subband obtained from multiresolution analysis of the speech signal. This method 1) calculates wavelet coefficients (subband information) up to the third level for each analysis frame length: 25 ms (frame shift: 10 ms), 2) simultaneously labels phonemes using phoneme segmentation, and 3) predicts high-frequency components using a pre-trained RNN-LSTM (Recurrent Neural Network with Long Short Term Memory) network for each phoneme label. In the RNN-LSTM training process, using the original 16-kHz speech signal, the wavelet coefficient containing the highest frequency detail coefficient was set as the teacher signal, and the lower level (narrowband) detail coefficient was set as the learning signal. To evaluate the effectiveness of the proposed method, in simulations, we measured the perceptual evaluation of listening speech quality (PESQ) value, an objective evaluation scale linked to subjective evaluation. We obtained an average PESQ value of 3.3, which can be considered fair-to-good quality (quality of speech score: excellent 5, good 4, fair 3, poor 2, bad 1).

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Bandwidth Extension of Speech Signals Using RNN-LSTM Network with Multiresolution Analysis

  • Seiji Hayashi

摘要

Public telephone systems are bandwidth-limited to 0.3 \(\sim \) 3.4 kHz due to telephone communication regulations, which significantly reduces voice quality and intelligibility. However, on the receiver side, it is possible to improve the clarity and intelligibility of the output speech quality by efficiently estimating and adding the high-frequency components lost from bandwidth limitation using only the narrow-band components. In this study, we use machine learning to estimate the high-frequency components in the range of 3.4 \(\sim \) 7 kHz that are normally lost due to bandwidth limitation, by utilizing the correlation between each subband obtained from multiresolution analysis of the speech signal. This method 1) calculates wavelet coefficients (subband information) up to the third level for each analysis frame length: 25 ms (frame shift: 10 ms), 2) simultaneously labels phonemes using phoneme segmentation, and 3) predicts high-frequency components using a pre-trained RNN-LSTM (Recurrent Neural Network with Long Short Term Memory) network for each phoneme label. In the RNN-LSTM training process, using the original 16-kHz speech signal, the wavelet coefficient containing the highest frequency detail coefficient was set as the teacher signal, and the lower level (narrowband) detail coefficient was set as the learning signal. To evaluate the effectiveness of the proposed method, in simulations, we measured the perceptual evaluation of listening speech quality (PESQ) value, an objective evaluation scale linked to subjective evaluation. We obtained an average PESQ value of 3.3, which can be considered fair-to-good quality (quality of speech score: excellent 5, good 4, fair 3, poor 2, bad 1).