<p>Bangla (Bengali) voice synthesis presents distinct issues due to its restricted language resources and complex vocabulary. This work introduces a state-of-the-art Bangla text-to-speech (TTS) system that utilizes the Variational Inference Synthesis of Speech (VITS) architecture. We tailor VITS for the Bengali language by utilizing a 20-hour custom-made dataset consisting of speech samples from a single speaker from native literary works. The model undergoes comprehensive training, enabling it to convert text into waveforms while detecting subtle pronunciation variations accurately. This study introduces a Bangla TTS system based on VITS, which has attained a Mean Opinion Score (MOS) of 4.01, a Mel-Cepstral Distortion (MCD) ranging from 5.24 to 5.48, a Word Error Rate (WER) between 0.16 and 0.18, and a Phoneme Error Rate (PER) of 0.15 to 0.17. These performance indicators validate the system’s excellence compared to current alternatives like Google Bangla TTS and Narakeet. Our model exhibits strong generalization capabilities, particularly with previously unseen complex vocabulary. The model has demonstrated superior performance to publicly accessible Bangla TTS models such as Google Bangla TTS, Narakeet, and AI4Bharat. Objective metrics validate the high level of comprehensibility and resemblance among speakers. Preliminary informal hearing tests indicate that the speech sounds highly natural. This study demonstrates the current advancements in neural text-to-speech (TTS) technology, which can be used in low-resource Indian languages, provided sufficient data is available.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Advancing Bangla text-to-speech synthesis using a VITS-based model with a custom dataset and comprehensive evaluation

  • Sujeet Kumar,
  • Siddharth Kumar,
  • Kushal Sathe,
  • Jayadeep Pati

摘要

Bangla (Bengali) voice synthesis presents distinct issues due to its restricted language resources and complex vocabulary. This work introduces a state-of-the-art Bangla text-to-speech (TTS) system that utilizes the Variational Inference Synthesis of Speech (VITS) architecture. We tailor VITS for the Bengali language by utilizing a 20-hour custom-made dataset consisting of speech samples from a single speaker from native literary works. The model undergoes comprehensive training, enabling it to convert text into waveforms while detecting subtle pronunciation variations accurately. This study introduces a Bangla TTS system based on VITS, which has attained a Mean Opinion Score (MOS) of 4.01, a Mel-Cepstral Distortion (MCD) ranging from 5.24 to 5.48, a Word Error Rate (WER) between 0.16 and 0.18, and a Phoneme Error Rate (PER) of 0.15 to 0.17. These performance indicators validate the system’s excellence compared to current alternatives like Google Bangla TTS and Narakeet. Our model exhibits strong generalization capabilities, particularly with previously unseen complex vocabulary. The model has demonstrated superior performance to publicly accessible Bangla TTS models such as Google Bangla TTS, Narakeet, and AI4Bharat. Objective metrics validate the high level of comprehensibility and resemblance among speakers. Preliminary informal hearing tests indicate that the speech sounds highly natural. This study demonstrates the current advancements in neural text-to-speech (TTS) technology, which can be used in low-resource Indian languages, provided sufficient data is available.