Low-Resource VITS-Based Emotion Speech Synthesis Using KNN Algorithm
摘要
Recently, the application scenarios of Text-to-Speech (TTS) have been expanding. Along with the simple replication of voice tones, there’s a growing demand for multi-emotional speech. However, traditional emotional speech synthesis typically relies on a large amount of annotated data, which might be scarce in certain applications or domains. In this paper, we aim to synthesize multi-emotional speech with minimal neutral samples using the K-Nearest Neighbors (KNN) algorithm and an enhanced VITS model for speech synthesis. Experimental results indicate that our approach improves both emotional expressiveness and speech clarity compared to traditional methods. Additionally, it achieves low-cost multi-emotional speech synthesis, making it suitable for resource-constrained applications.