Expressive Speech Synthesis Enhancement with Conditional Embeddings
摘要
In recent years, emotional speech synthesis techniques have attracted considerable interest because of their wide-ranging potential applications. However, when confronted with datasets containing emotional attributes, speech synthesized by traditional methods frequently encounters difficulties, such as a mismatch with the text content or an unnatural expression of emotions. To solve these problems, we developed a straightforward emotional speech synthesis model. This model builds upon the VITS framework and incorporates an emotion prediction module, a prosody prediction module, and a conditional encoder. It can automatically predict emotion labels based on the input text, or manually specify emotions for precise control of emotions. Experimental outcomes indicate improved naturalness and expressiveness in synthesized speech, thus enhancing the overall audio quality.