Grapheme-to-Phoneme Conversion for Cyrillic Mongolian Using a Speech Corpus
摘要
In this research work, we analyzed cases of different pronunciations for the same written text (homographs) and aimed to identify sounds pronounced but not written by creating a phoneme-aligned speech corpus for Cyrillic Mongolian script. The speech corpus consists of 25,791 sentences designed for text-to-speech conversion, containing 60.28 h of studio recordings from a single speaker, along with corresponding transcripts. Utilizing the phoneme alignments generated through this research, we developed a transformer-based grapheme-to-phoneme converter. This model was trained on the phoneme corpus, leveraging the rich phonetic information extracted from the aligned speech data. We then conducted a baseline evaluation of the grapheme-to-phoneme converter to assess its performance and accuracy. This comprehensive approach, combining corpus creation, phoneme alignment analysis, and machine learning model development, provides valuable insights into pronunciation variations and advances the field of text-to-speech technology. The results of this study contribute to a better understanding of the relationship between written text and spoken language, while also establishing a foundation for future improvements in grapheme-to-phoneme conversion techniques for Cyrillic Mongolian script.