An Efficient Emotional Speech Synthesis Approach via Multi-scale Feature Generation
摘要
In recent years, emotional speech synthesis technology has advanced considerably, with substantial improvements in both synthesis quality and emotional expression. However, current methods and models are either too complex with cumbersome processing workflows, making them impractical for real-time applications, or they generate speech with relatively rough emotions, lacking the nuanced expressiveness needed for more refined emotional delivery. To tackle these challenges, we propose an emotional feature generation method with multi-feature fusion capabilities. Based on this foundation, we further introduce an emotional speech synthesis approach using the ProDiff model. This method enhances speech emotional expressiveness by predicting and generating emotional features across multiple scales, thereby improving the quality of emotional speech synthesis. Finally, the improved model is assessed to validate the effectiveness of the proposed enhancements.