Automatic speech recognition (ASR) for Kazakh presents notable challenges primarily due to the language’s agglutinative characteristics and intricate morphological variations. The article describes the process of collecting and processing audio data in the Kazakh language for speech recognition. Whisper OpenAI, Soyle, and Voiser models were used for speech-to-text recognition, and experiments on Kazakh speech recognition were conducted using them. Speech recognition was evaluated using the BLEU, WER, and TER metrics. Among the models, Soyle and Voiser demonstrated the highest BLEU scores of 64% and 63%, while the Whisper model only reached 20.9%. The lower WER and TER scores of 27% and 24% of the Soyle and Voiser models also showed their supremacy on the Whisper model, where TER and WER scores were 57%. Generally, all models improved the quality of transcriptions, making them more suitable for practical use. This emphasizes the critical need for advancing ASR technologies tailored to the Kazakh language and suggests new directions for future research aimed at finding an optimal balance between model complexity and efficiency. Consequently, this study contributes meaningfully to enhancing speech recognition technologies for Kazakh, paving the way for more effective applications in the field.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Collecting, Processing, and Evaluating the Performance of Kazakh Automatic Speech Recognition

  • Aidana Karibayeva,
  • Vladislav Karyukin,
  • Diana Rakhimova,
  • Dina Amirova,
  • Rashid Aliyev,
  • Adina Karybayeva

摘要

Automatic speech recognition (ASR) for Kazakh presents notable challenges primarily due to the language’s agglutinative characteristics and intricate morphological variations. The article describes the process of collecting and processing audio data in the Kazakh language for speech recognition. Whisper OpenAI, Soyle, and Voiser models were used for speech-to-text recognition, and experiments on Kazakh speech recognition were conducted using them. Speech recognition was evaluated using the BLEU, WER, and TER metrics. Among the models, Soyle and Voiser demonstrated the highest BLEU scores of 64% and 63%, while the Whisper model only reached 20.9%. The lower WER and TER scores of 27% and 24% of the Soyle and Voiser models also showed their supremacy on the Whisper model, where TER and WER scores were 57%. Generally, all models improved the quality of transcriptions, making them more suitable for practical use. This emphasizes the critical need for advancing ASR technologies tailored to the Kazakh language and suggests new directions for future research aimed at finding an optimal balance between model complexity and efficiency. Consequently, this study contributes meaningfully to enhancing speech recognition technologies for Kazakh, paving the way for more effective applications in the field.