This paper presents an investigation of the performance of various end-to-end ASR models trained on low-rescourced Livvi-Karelian. Several Wav2Vec 2.0 and Whisper based models were fine-tuned, tested and compared with the hybrid TDNN-F/HMM. In the course of the experiments, end-to-end Transformer-based models have demonstrated a good performance, however the best results obtained were due to a combination of N-gram and Transformer-based models. The result of 19.83% WER on the test set were obtained using the Wav2Vec 2.0 large model with N-gram augmentation, thus being on par with SOTA models for other low-resource languages. Besides, this paper presents a new language corpus of Livvi-Karelian, containing transcripts from radio broadcasts, featuring samples from 17 speakers (7 males and 10 females). Covering about 4.5 h of audio recordings, it contains 32,037 words, thus being a valuable tool for linguistic research. The findings of the presented work may be of considerable interest both for low-resource ASR and field Finnougristics.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Towards a Livvi-Karelian End-to-End ASR System

  • Irina Kipyatkova,
  • Ildar Kagirov,
  • Mikhail Dolgushin,
  • Alexandra Rodionova

摘要

This paper presents an investigation of the performance of various end-to-end ASR models trained on low-rescourced Livvi-Karelian. Several Wav2Vec 2.0 and Whisper based models were fine-tuned, tested and compared with the hybrid TDNN-F/HMM. In the course of the experiments, end-to-end Transformer-based models have demonstrated a good performance, however the best results obtained were due to a combination of N-gram and Transformer-based models. The result of 19.83% WER on the test set were obtained using the Wav2Vec 2.0 large model with N-gram augmentation, thus being on par with SOTA models for other low-resource languages. Besides, this paper presents a new language corpus of Livvi-Karelian, containing transcripts from radio broadcasts, featuring samples from 17 speakers (7 males and 10 females). Covering about 4.5 h of audio recordings, it contains 32,037 words, thus being a valuable tool for linguistic research. The findings of the presented work may be of considerable interest both for low-resource ASR and field Finnougristics.