Enhancing Grapheme-to-Phoneme Conversion for Bengali Speech Recognition: Leveraging Dialectical Variations and Morphophonemic Rules
摘要
The current work presents a Grapheme-to-Phoneme (G2P) conversion method that has been specifically engineered to accommodate the dialectal variations within the Bengali language. In contrast to previous studies, our contribution entails the integration of numerous Bengali dialects—Sylheti, Dhakaiya, Chittagong, Rarhi, and Bangali—into a unified G2P system. We utilize an attention-based sequence-to-sequence (Seq2Seq) neural network to capture the phonetic and morphophonemic variances of these dialects. The suggested model demonstrates enhanced transcription accuracy, achieving a Word Error Rate (WER) of 3.7% and a Phoneme Error Rate (PER) of 2.9%. In comparison, the baseline rates for WER and PER are 8.5% and 6.3% respectively. The model's out-of-vocabulary (OOV) word accuracy also improved from 70% to 85%. This study delves deeper into the model's ability to handle larger workloads in real-time applications and its potential to be adjusted for use in other languages that lack sufficient resources. This goes beyond being merely a theoretical contribution.