Collecting, Processing, and Evaluating the Performance of Kazakh Automatic Speech Recognition
摘要
Automatic speech recognition (ASR) for Kazakh presents notable challenges primarily due to the language’s agglutinative characteristics and intricate morphological variations. The article describes the process of collecting and processing audio data in the Kazakh language for speech recognition. Whisper OpenAI, Soyle, and Voiser models were used for speech-to-text recognition, and experiments on Kazakh speech recognition were conducted using them. Speech recognition was evaluated using the BLEU, WER, and TER metrics. Among the models, Soyle and Voiser demonstrated the highest BLEU scores of 64% and 63%, while the Whisper model only reached 20.9%. The lower WER and TER scores of 27% and 24% of the Soyle and Voiser models also showed their supremacy on the Whisper model, where TER and WER scores were 57%. Generally, all models improved the quality of transcriptions, making them more suitable for practical use. This emphasizes the critical need for advancing ASR technologies tailored to the Kazakh language and suggests new directions for future research aimed at finding an optimal balance between model complexity and efficiency. Consequently, this study contributes meaningfully to enhancing speech recognition technologies for Kazakh, paving the way for more effective applications in the field.