PMIVC: Perturbation with Mutual Information Learning for One-Shot Voice Conversion
摘要
One-shot Voice Conversion (VC) is distinguished by its capability to transform a speaker’s identity using just a single target voice sample. Current methods typically decouple acoustic features to achieve this transformation. However, the intricate interactions between these features often result in the entanglement of timbre with content and prosody, leading to issues of timbre leakage in the converted audio. In this paper, we propose a novel method called Perturbation with Mutual Information Learning for One-shot Voice Conversion (PMIVC), leveraging recent advances in information perturbation. Initially, the voice signal is fed into two different perturbation modules to eliminate redundant information in the input features corresponding to the content and prosody encoders. Subsequently, the processed voice signal is input into various encoders, where mutual information learning is employed to further decouple the acoustic features and reduce their correlation with timbre characteristics. The final converted audio is then generated through a decoder and a vocoder. Our experimental results clearly confirm that the PMIVC method we proposed excels in enhancing the natural flow of speech and preserving the uniqueness of the speaker’s characteristics, achieving significant progress and showing a clear advantage over traditional models. In addition, we can effectively decouple acoustic features and reduce the risk of speaker timbre leakage.