USIAL-VC: A One-Shot Voice Conversion by U-Net-Based Encoder and Speaker Identity Adaptive Learning
摘要
Voice conversion (VC) is an audio processing technology that converts the source voice into the target voice of another speaker without changing the linguistic content. There is still a significant gap between the target and converted voices in terms of voice quality and speaker similarity in one-shot VC, which remains a challenging issue to address. Feature disentanglement models have been previously used to separate speaker and audio content information. However, achieving effective disentanglement is challenging, which limits the practical applicability of these models. This paper proposes USIAL-VC, an improved U-net-based encoder and Speaker Identity Adaptive Learning model. The encoder of USIAL-VC employs a U-net architecture for down-sampling. The bottleneck layer of the model combines instance normalization and vector quantization to filter out speaker identity. Additionally, the residual information following the bottleneck layer is leveraged to adaptively learn the true speaker identity. Both objective and subjective results demonstrate that the proposed approach effectively captures the characteristics of the target speaker while preserving audio quality. Furthermore, in one-shot VC, the proposed method maintains strong performance in both audio quality and speaker similarity compared to other state-of-the-art VC models, even with only 30 s of target speech.