Operatic Singing Voice Synthesis From Inexperienced Voice Considering Tempo and Vowel Change
摘要
The purpose of this study is to develop a system that can synthesize operatic singing voices from the speaking voice of a user who has no experience singing opera. In our previous work, a method that converts the speaker individuality of the input opera singing voice into that of the target user by using Diff-SVC-based voice conversion has been proposed. However, the conventional system requires a professional operatic singing voice as an input for voice conversion, making it difficult to synthesize arbitrary songs. In this paper, we propose using operatic singing voices synthesized from scores using DiffSinger instead of using actual operatic singing voices as the input to a Diff-SVC-based voice conversion system for arbitrary songs synthesis. Our study found that the conventional DiffSinger produces singing voices with less “Operatic-ness” compared to actual professional opera singing. To solve this problem, we propose to introduce a tempo estimator and a vowel change estimator into DiffSinger of the above user opera singing synthesis system. The tempo estimator and the vowel change estimator estimate tempo changes and the pronunciation tendency of vowels in a cappella opera singing from the input score features and embed them into the score features. By using these score features for singing voice synthesis, it is possible to synthesize operatic singing voices that take the characteristics of a cappella opera singing more into account. The effectiveness of proposed methods was confirmed through 5-stage MOS evaluation experiments, an A/B test experiment and objective evaluation experiments.