Robust speech recognition method based on dense time–frequency convolution and bispectral refinement enhancement
摘要
In noisy environments, the accuracy of speech recognition is often affected by noise interference, making it necessary to use speech enhancement techniques to mitigate this impact. Current methods that incorporate enhancement modules through joint training have made some progress in improving the robustness of speech recognition systems. However, these approaches primarily focus on directly modeling the original noisy signal without fully considering the use of potentially valuable information within the noise to support the enhancement modeling process. To address these issues, this work proposes a robust speech recognition method based on dense time–frequency convolution and bispectral refinement enhancement. This method attempts to avoid the problem of excessive suppression by repairing distortions that may be introduced by the speech enhancement module while performing noise reduction. First, a single-channel speech enhancement module based on dense time–frequency convolution is used for initial noise suppression. Then, a bispectral refinement enhancement module is designed to extract beneficial features from the estimated noise to improve speech quality. Finally, a proposed weighted speech distortion loss function is applied through multi-task joint training to further enhance recognition performance. Experimental results show that the proposed method reduces the word error rate in speech recognition by 17.57% compared to baseline methods.