DynFusion-VC: Zero-Shot Voice Conversion via Optimized Fusion of Features
摘要
Zero-shot voice conversion (VC) aims to transform a source speaker’s voice into that of a target speaker, without requiring any target speaker data during training. However, existing methods are hard to disentangle speech content from speaker-specific features, leading to speaker information leakage. Additionally, these methods often involve high computational costs for both training and inference. In this paper, we propose DynFusion-VC, a novel zero-shot voice conversion method that achieves high-quality results with low computational complexity. DynFusion-VC is text-free, any-to-any, and reduces computational overhead through a series of innovative techniques. Specifically, we leverage data augmentation to generate additional target speech for better matching, employ a self-supervised model to extract speaker-independent features from each speech frame, and apply a dynamic matching strategy to synthesize the target features for the converted speech. Finally, we generate the converted speech using a vocoder. Experimental results, including both objective and subjective evaluations, demonstrate that DynFusion-VC outperforms state-of-the-art VC models in naturalness, intelligibility, and speaker similarity.