Speech synthesis is striving to explore efficient content representations for different tasks, which also drives research on models capable of performing multiple tasks. However, current approaches are predominantly based on text-to-speech models, which require large-scale transcribed datasets, simultaneously limiting the representations. To address this issue, we propose EmbSpeech, a novel framework unifying zero-shot speech synthesis tasks through speaker-independent self-supervised speech representations. EmbSpeech is built on a voice conversion model and consists of three task-specific encoders, a speaker encoder, and a diffusion-based decoder. With task-specific encoders, EmbSpeech can perform voice conversion, text-to-speech synthesis, and speaking style conversion. Experiments show that our framework can achieve naturalness and intelligibility comparable to or even better than the baselines in low-resource settings.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

EmbSpeech: A Unified Framework Towards Low-Resource Zero-Shot Speech Synthesis

  • Fan Huang,
  • Dong Chen,
  • Chengxin Chen,
  • Zhengxuan Song,
  • Ge Lin,
  • Kun Zeng

摘要

Speech synthesis is striving to explore efficient content representations for different tasks, which also drives research on models capable of performing multiple tasks. However, current approaches are predominantly based on text-to-speech models, which require large-scale transcribed datasets, simultaneously limiting the representations. To address this issue, we propose EmbSpeech, a novel framework unifying zero-shot speech synthesis tasks through speaker-independent self-supervised speech representations. EmbSpeech is built on a voice conversion model and consists of three task-specific encoders, a speaker encoder, and a diffusion-based decoder. With task-specific encoders, EmbSpeech can perform voice conversion, text-to-speech synthesis, and speaking style conversion. Experiments show that our framework can achieve naturalness and intelligibility comparable to or even better than the baselines in low-resource settings.