Recent advancements in end-to-end Text-to-Speech (TTS) systems have significantly improved speech naturalness. However, existing models often struggle with speaker adaptation in low-resource settings, leading to degraded speaker characteristics, unnatural prosody, and pronunciation errors, particularly in languages with complex phonetic structures. Vietnamese, a tonal language with six distinct tones, intricate pronunciation rules, and diverse regional accents, presents unique challenges for TTS synthesis. In this study, we introduce VIEACT-TTS-VC, a neural TTS and voice conversion system based on the VITS framework, designed to synthesize natural and intelligible Vietnamese speech while effectively adapting to speaker styles with limited data. The proposed system offers three key functionalities: 1) generating high-quality speech with style adaptation across diverse speakers, 2) enabling robust adaptation to new speech domains using small datasets, and 3) functioning as a hybrid system capable of both text-to-speech synthesis and voice conversion while preserving linguistic content. To evaluate the performance of VIEACT-TTS-VC, we employ both subjective and objective assessment metrics. Subjective evaluations, conducted through listener-based Mean Opinion Score, yield 4.31 ± 0.158 for comprehensibility and 3.68 ± 0.135 for naturalness with only 5 min of additional training data. Objective evaluations further demonstrate that VIEACT-TTS-VC surpasses state-of-the-art models in speaker style adaptation and exhibits remarkable robustness in low-resource scenarios.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

VIEACT-TTS-VC: A Vietnamese End-to-End Text-to-Speech and Voice Conversion Framework Using Low-Resource Adaptable Style Transfer

  • Hung D. Vo,
  • Long S. T. Nguyen,
  • Tri H. Trinh,
  • Khang H. N. Vo,
  • Tho T. Quan

摘要

Recent advancements in end-to-end Text-to-Speech (TTS) systems have significantly improved speech naturalness. However, existing models often struggle with speaker adaptation in low-resource settings, leading to degraded speaker characteristics, unnatural prosody, and pronunciation errors, particularly in languages with complex phonetic structures. Vietnamese, a tonal language with six distinct tones, intricate pronunciation rules, and diverse regional accents, presents unique challenges for TTS synthesis. In this study, we introduce VIEACT-TTS-VC, a neural TTS and voice conversion system based on the VITS framework, designed to synthesize natural and intelligible Vietnamese speech while effectively adapting to speaker styles with limited data. The proposed system offers three key functionalities: 1) generating high-quality speech with style adaptation across diverse speakers, 2) enabling robust adaptation to new speech domains using small datasets, and 3) functioning as a hybrid system capable of both text-to-speech synthesis and voice conversion while preserving linguistic content. To evaluate the performance of VIEACT-TTS-VC, we employ both subjective and objective assessment metrics. Subjective evaluations, conducted through listener-based Mean Opinion Score, yield 4.31 ± 0.158 for comprehensibility and 3.68 ± 0.135 for naturalness with only 5 min of additional training data. Objective evaluations further demonstrate that VIEACT-TTS-VC surpasses state-of-the-art models in speaker style adaptation and exhibits remarkable robustness in low-resource scenarios.