Automatic speech recognition (ASR) technologies have advanced significantly in the past years; yet, the creation of effective ASR models for low-resource languages remains a challenge. This research describes an approach that seeks to improve the transcription quality of Gujarati, a low-resource Indian language, in terms of fine-tuning the Wav2vec2-XLSR model. By training on two open-source datasets, OpenSLR and Kathbath, the model is optimized for Gujarati’s unique linguistic features. This approach demonstrates the model’s ability to adapt the phonetic complexities of Gujarati Language. Our methodology consisted fine-tuning on two diverse corpus with significant increase in transcription accuracy for Gujarati language despite the data limitations. The results of this study highlights the potential of self-supervised learning models to transform ASR applications for low-resource languages, making transcription solutions more accessible. This research not only contributes to ASR for Gujarati but also provides a model for enhancing ASR systems in various other resource-constrained environments. This research underscores promising directions for future advancements in ASR for low resource languages.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimizing ASR for Low-Resource Language: Fine-Tuning Wav2Vec2-XLSR for Gujarati

  • Drashti Joshi,
  • Taneeshk Patel,
  • Zalak Kansagra

摘要

Automatic speech recognition (ASR) technologies have advanced significantly in the past years; yet, the creation of effective ASR models for low-resource languages remains a challenge. This research describes an approach that seeks to improve the transcription quality of Gujarati, a low-resource Indian language, in terms of fine-tuning the Wav2vec2-XLSR model. By training on two open-source datasets, OpenSLR and Kathbath, the model is optimized for Gujarati’s unique linguistic features. This approach demonstrates the model’s ability to adapt the phonetic complexities of Gujarati Language. Our methodology consisted fine-tuning on two diverse corpus with significant increase in transcription accuracy for Gujarati language despite the data limitations. The results of this study highlights the potential of self-supervised learning models to transform ASR applications for low-resource languages, making transcription solutions more accessible. This research not only contributes to ASR for Gujarati but also provides a model for enhancing ASR systems in various other resource-constrained environments. This research underscores promising directions for future advancements in ASR for low resource languages.