Comparison of Well and Lower-Resourced Self-training in ASR
摘要
In this paper, we present a case study of self-training for end-to-end Hungarian and Mandarin speech recognition. We demonstrate that self-training methods can significantly improve the accuracy of the baseline model for both languages. We trained a supervised baseline model which achieved a 20.13% WER on the BEA-Base eval-spont set. The self-trained Hungarian model, combining pseudo-labels generated by the baseline seed model with labeled data, achieved a WER of 15.93% on the BEA-Base eval-spont set, representing a 20.84% relative reduction compared to the baseline, and reduced WER by relative 15.28% on the independent Common Voice test set. The Mandarin model, which relied solely on pseudo-labels from the Whisper large-V2 model and used no labeled data, reduced CER by 45.31% on the AISHELL-2018A-EVAL test set, improving from 13.00% to 7.11%, and by relative 32.42% on the external Common Voice test set. We find that for spontaneous, low-resource Hungarian ASR tasks, pseudo-labels from domain-specific models are more effective than those from large general models like Whisper large-V2.