Zero-Shot Voice Cloning Based on Target Adaptation and Phoneme-Level Local Feature Embedding
摘要
Voice cloning is the process of creating a digital facsimile of a target speaker’s voice from minimal samples. Contemporary TTS models proficiently generate speech that mirrors the voices of familiar speakers. However, cloning the voice of an unseen speaker with no prior data presents a greater challenge. In this paper, we present a novel zero-shot voice cloning method designed to enhance the generalization to unseen speakers while simultaneously improving the prosody of the generated speech. Our approach utilizes the ECAPA-TDNN framework as a speaker encoder for distilling global speaker characteristics. We introduce a target adaptation layer normalization technique, which incorporates a unique target adaptation encoder. This encoder derives a specific adaptation vector from the reference speech to facilitate data normalization, thereby significantly improving the model’s ability to generalize. Moreover, the model is augmented with dual phoneme-level attention modules that capture nuanced phonetic embeddings from the reference audio, thereby enriching the prosodic expressiveness and fidelity to the target speaker’s voice. Empirical evaluations reveal that our model achieves high fidelity in voice synthesis, particularly excelling in zero-shot scenarios involving unseen speakers, and marks a substantial leap forward in the quest for more adaptable and expressive TTS systems.