<p>In noisy environments, the accuracy of speech recognition is often affected by noise interference, making it necessary to use speech enhancement techniques to mitigate this impact. Current methods that incorporate enhancement modules through joint training have made some progress in improving the robustness of speech recognition systems. However, these approaches primarily focus on directly modeling the original noisy signal without fully considering the use of potentially valuable information within the noise to support the enhancement modeling process. To address these issues, this work proposes a robust speech recognition method based on dense time–frequency convolution and bispectral refinement enhancement. This method attempts to avoid the problem of excessive suppression by repairing distortions that may be introduced by the speech enhancement module while performing noise reduction. First, a single-channel speech enhancement module based on dense time–frequency convolution is used for initial noise suppression. Then, a bispectral refinement enhancement module is designed to extract beneficial features from the estimated noise to improve speech quality. Finally, a proposed weighted speech distortion loss function is applied through multi-task joint training to further enhance recognition performance. Experimental results show that the proposed method reduces the word error rate in speech recognition by 17.57% compared to baseline methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Robust speech recognition method based on dense time–frequency convolution and bispectral refinement enhancement

  • Wenjun Wang,
  • Ling Dong,
  • Zhengtao Yu,
  • Yuxin Huang,
  • Shangbin Mo,
  • Linqing Wang

摘要

In noisy environments, the accuracy of speech recognition is often affected by noise interference, making it necessary to use speech enhancement techniques to mitigate this impact. Current methods that incorporate enhancement modules through joint training have made some progress in improving the robustness of speech recognition systems. However, these approaches primarily focus on directly modeling the original noisy signal without fully considering the use of potentially valuable information within the noise to support the enhancement modeling process. To address these issues, this work proposes a robust speech recognition method based on dense time–frequency convolution and bispectral refinement enhancement. This method attempts to avoid the problem of excessive suppression by repairing distortions that may be introduced by the speech enhancement module while performing noise reduction. First, a single-channel speech enhancement module based on dense time–frequency convolution is used for initial noise suppression. Then, a bispectral refinement enhancement module is designed to extract beneficial features from the estimated noise to improve speech quality. Finally, a proposed weighted speech distortion loss function is applied through multi-task joint training to further enhance recognition performance. Experimental results show that the proposed method reduces the word error rate in speech recognition by 17.57% compared to baseline methods.