Realizing Natural Prosody in Statistical Speech Synthesis
摘要
High-quality synthetic speech is now available by waveform concatenation. However, the quality is supported by a huge speech corpus, and it is difficult to realize speech not included in the corpus. In order to increase “flexibility” of speech synthesis, methods handling speech as acoustic parameters need to be improved. HMM-based speech synthesis handles acoustic parameters in statistical way and can generate speech with new voice qualities/utterance styles only from a limited speech corpus. However, it has an inherent problem for prosody; handling acoustic parameters without a viewpoint of wider time-span. Prosodic features are related to words, phrases, sentences, and even paragraphs. Generation process model of fundamental frequency (F0) contours is ideal to represent global features of prosody. A method is introduced which decomposes F0 contours into three layers, and handles them differently in the speech synthesis process. Also issues of flexible control of prosody are addressed.