In-vehicle automatic speech recognition plays a crucial role in the field of autonomous driving and in-car voice assistants, with one of the significant factors affecting recognition accuracy being noise interference in the vehicle environment. Most advanced automatic speech recognition methods are oriented towards training resource-rich languages and have some limitations for other small languages. This paper proposes a Collaborative Transformer Decoder Method (CTDM) for a Low-resource Uyghur Speech Recognition in-vehicle Environment. In the encoder, we adopt collaborative encoding to capture multi-scale detail information while focusing on global information. In the decoder part, we design a parallel decoding strategy for arbitrary sequences, breaking the traditional left-to-right or right-to-left decoding order in Transformer decoding methods to establish associations between different characters. Our CTDM enhances the model’s ability to extract detailed information, reduces the model’s dependence on large-scale training data, and weakens the impact of noise on the model. Experiments are conducted on the Uyghur General Speech 7, 8, 9, and 16 datasets combined with vehicle noise and human speech noise to simulate real vehicle speech environments. Experimental results indicate that in a vehicular noise environment with a signal-to-noise ratio (SNR) of 0, the average word error rates (WER) on datasets 7, 8, 9, and 16 decreased by 29.8%, 15.9%, 8.7%, and 5.1%, respectively. In a vocal noise environment with an SNR of 0, the WER on datasets 7, 8, 9, and 16 decreased by 11.0%, 35.5%, 32.9%, and 13.8%, respectively.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Collaborative Transformer Decoder Method for Uyghur Speech Recognition in-Vehicle Environment

  • Jiang Zhang,
  • Liejun Wang,
  • Yinfeng Yu,
  • Miaomiao Xu,
  • Alimjan Mattursun

摘要

In-vehicle automatic speech recognition plays a crucial role in the field of autonomous driving and in-car voice assistants, with one of the significant factors affecting recognition accuracy being noise interference in the vehicle environment. Most advanced automatic speech recognition methods are oriented towards training resource-rich languages and have some limitations for other small languages. This paper proposes a Collaborative Transformer Decoder Method (CTDM) for a Low-resource Uyghur Speech Recognition in-vehicle Environment. In the encoder, we adopt collaborative encoding to capture multi-scale detail information while focusing on global information. In the decoder part, we design a parallel decoding strategy for arbitrary sequences, breaking the traditional left-to-right or right-to-left decoding order in Transformer decoding methods to establish associations between different characters. Our CTDM enhances the model’s ability to extract detailed information, reduces the model’s dependence on large-scale training data, and weakens the impact of noise on the model. Experiments are conducted on the Uyghur General Speech 7, 8, 9, and 16 datasets combined with vehicle noise and human speech noise to simulate real vehicle speech environments. Experimental results indicate that in a vehicular noise environment with a signal-to-noise ratio (SNR) of 0, the average word error rates (WER) on datasets 7, 8, 9, and 16 decreased by 29.8%, 15.9%, 8.7%, and 5.1%, respectively. In a vocal noise environment with an SNR of 0, the WER on datasets 7, 8, 9, and 16 decreased by 11.0%, 35.5%, 32.9%, and 13.8%, respectively.